{"id":{"repo_id":"york","oai_identifier":"oai:yorkspace.library.yorku.ca:10315/42384"},"canonical_url":"https://search.dev.ndltd.org/etd/york/oai:yorkspace.library.yorku.ca:10315/42384","repository":{"repo_id":"york","name":"York University","base_url":"https://yorkspace.library.yorku.ca/oai/request"},"display":{"title":"Studying The Effectiveness Of Large Language Models In Benchmark Biomedical Tasks","abstract":"Recently, Large Language Models (LLMs) have demonstrated impressive capability to solve a wide range of tasks. However, despite their success across various tasks, no prior work has investigated their capability in the biomedical domain yet. To this end, this thesis aims to evaluate the performance of LLMs on benchmark biomedical tasks. For this purpose, a comprehensive evaluation of 4 popular LLMs in 6 diverse biomedical tasks across 26 datasets has been conducted. Interestingly, this evaluation shows that in biomedical datasets that have smaller training sets, zero-shot LLMs even outperform the current state-of-the-art models when they were fine-tuned only on the training set of these datasets. This suggests that pretraining on large text corpora makes LLMs quite specialized even in the biomedical domain. The findings also shows that not a single LLM can outperform other LLMs in all tasks, with the performance of different LLMs may vary depending on the task. While their performance is still quite poor in comparison to the biomedical models that were fine-tuned on large training sets, this study demonstrates that LLMs have the potential to be a valuable tool for various biomedical tasks that lack large annotated data.","abstract_html":"Recently, Large Language Models (LLMs) have demonstrated impressive capability to solve a wide range of tasks. However, despite their success across various tasks, no prior work has investigated their capability in the biomedical domain yet. To this end, this thesis aims to evaluate the performance of LLMs on benchmark biomedical tasks. For this purpose, a comprehensive evaluation of 4 popular LLMs in 6 diverse biomedical tasks across 26 datasets has been conducted. Interestingly, this evaluation shows that in biomedical datasets that have smaller training sets, zero-shot LLMs even outperform the current state-of-the-art models when they were fine-tuned only on the training set of these datasets. This suggests that pretraining on large text corpora makes LLMs quite specialized even in the biomedical domain. The findings also shows that not a single LLM can outperform other LLMs in all tasks, with the performance of different LLMs may vary depending on the task. While their performance is still quite poor in comparison to the biomedical models that were fine-tuned on large training sets, this study demonstrates that LLMs have the potential to be a valuable tool for various biomedical tasks that lack large annotated data.","abstract_has_math":false,"creators":["Jahan, Israt"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Huang, Jimmy Xiangji"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-10-28","date_published":"2024-10-28","updated_at":"2026-07-24T06:33:40Z","subjects":["Biology","Artificial intelligence","Bioinformatics"],"languages":["en"],"rights":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10315/42384","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Huang, Jimmy Xiangji"]},{"key":"dc:creator","label":"Author","values":["Jahan, Israt"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-10-28T13:37:20Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-10-28T13:37:20Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-10-28"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Biology","Artificial intelligence","Bioinformatics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10315/42384"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Recently, Large Language Models (LLMs) have demonstrated impressive capability to solve a wide range of tasks. However, despite their success across various tasks, no prior work has investigated their capability in the biomedical domain yet. To this end, this thesis aims to evaluate the performance of LLMs on benchmark biomedical tasks. For this purpose, a comprehensive evaluation of 4 popular LLMs in 6 diverse biomedical tasks across 26 datasets has been conducted. Interestingly, this evaluation shows that in biomedical datasets that have smaller training sets, zero-shot LLMs even outperform the current state-of-the-art models when they were fine-tuned only on the training set of these datasets. This suggests that pretraining on large text corpora makes LLMs quite specialized even in the biomedical domain. The findings also shows that not a single LLM can outperform other LLMs in all tasks, with the performance of different LLMs may vary depending on the task. While their performance is still quite poor in comparison to the biomedical models that were fine-tuned on large training sets, this study demonstrates that LLMs have the potential to be a valuable tool for various biomedical tasks that lack large annotated data."]},{"key":"dc:title","label":"Title","values":["Studying The Effectiveness Of Large Language Models In Benchmark Biomedical Tasks"]}]}],"canonical_facts":{"dc:contributor.advisor":["Huang, Jimmy Xiangji"],"dc:creator":["Jahan, Israt"],"dc:date.accessioned":["2024-10-28T13:37:20Z"],"dc:date.available":["2024-10-28T13:37:20Z"],"dc:date.issued":["2024-10-28"],"dc:description.abstract":["Recently, Large Language Models (LLMs) have demonstrated impressive capability to solve a wide range of tasks. However, despite their success across various tasks, no prior work has investigated their capability in the biomedical domain yet. To this end, this thesis aims to evaluate the performance of LLMs on benchmark biomedical tasks. For this purpose, a comprehensive evaluation of 4 popular LLMs in 6 diverse biomedical tasks across 26 datasets has been conducted. Interestingly, this evaluation shows that in biomedical datasets that have smaller training sets, zero-shot LLMs even outperform the current state-of-the-art models when they were fine-tuned only on the training set of these datasets. This suggests that pretraining on large text corpora makes LLMs quite specialized even in the biomedical domain. The findings also shows that not a single LLM can outperform other LLMs in all tasks, with the performance of different LLMs may vary depending on the task. While their performance is still quite poor in comparison to the biomedical models that were fine-tuned on large training sets, this study demonstrates that LLMs have the potential to be a valuable tool for various biomedical tasks that lack large annotated data."],"dc:identifier.uri":["https://hdl.handle.net/10315/42384"],"dc:language":["en"],"dc:rights":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."],"dc:subject":["Biology","Artificial intelligence","Bioinformatics"],"dc:title":["Studying The Effectiveness Of Large Language Models In Benchmark Biomedical Tasks"],"dc:type":["Electronic Thesis or Dissertation"]},"updated_at":"2026-07-24T06:33:40Z"}