{"id":{"repo_id":"regina","oai_identifier":"oai:uregina.scholaris.ca:10294/15008"},"canonical_url":"https://search.dev.ndltd.org/etd/regina/oai:uregina.scholaris.ca:10294/15008","repository":{"repo_id":"regina","name":"University of Regina","base_url":"https://uregina.scholaris.ca/server/oai/request"},"display":{"title":"Automatic Summarization of Financial Reports","abstract":"The field of Natural Language Processing (NLP) has witnessed substantial advancements due to both the development of new algorithms and the increase of computational capacities. As a result, novel NLP techniques have been successfully applied in various fields such as medicine, law, and finance. Companies and corporations publish several financial reports and disclosures every year to meet the requirements and needs of law enforcement entities, analysts, customers, and stakeholders. Processing these materials requires a tremendous amount of time and manual labor, especially with the unwieldy volume of data accumulated throughout the years. This thesis attempts to exploit the potential of modern NLP techniques to summarize financial reports in an automated fashion. We utilize an attention-based language model, namely RoBERTa to build a financially-oriented language model. This model is combined with various non-parametric methods to generate sentence vectors which can facilitate the efficient selection of important sentences in the document as part of the summarization algorithm. We illustrate the superiority of a financially tuned language model. Further, it is shown that simple non-parametric sentence embeddings achieve ROUGE scores that are comparable to or even better than those produced by Sentence Transformer and Universal Embeddings.","abstract_html":"The field of Natural Language Processing (NLP) has witnessed substantial advancements due to both the development of new algorithms and the increase of computational capacities. As a result, novel NLP techniques have been successfully applied in various fields such as medicine, law, and finance. Companies and corporations publish several financial reports and disclosures every year to meet the requirements and needs of law enforcement entities, analysts, customers, and stakeholders. Processing these materials requires a tremendous amount of time and manual labor, especially with the unwieldy volume of data accumulated throughout the years. This thesis attempts to exploit the potential of modern NLP techniques to summarize financial reports in an automated fashion. We utilize an attention-based language model, namely RoBERTa to build a financially-oriented language model. This model is combined with various non-parametric methods to generate sentence vectors which can facilitate the efficient selection of important sentences in the document as part of the summarization algorithm. We illustrate the superiority of a financially tuned language model. Further, it is shown that simple non-parametric sentence embeddings achieve ROUGE scores that are comparable to or even better than those produced by Sentence Transformer and Universal Embeddings.","abstract_has_math":false,"creators":["Alikhani, Majid"],"institution":"Faculty of Graduate Studies and Research, University of Regina","degree_name":"Master of Science (MSc)","degree_level":"Master&apos;s","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Zilles, Sandra","Sadeqi, Mehdi"],"committee_chairs":[],"committee_members":["Mouhoub, Malek"],"year":2021,"date_issued":"2021-09","date_published":"2021-09","updated_at":"2026-07-24T04:03:39Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.82465/4551"],"render_values":[{"text":"https://doi.org/10.82465/4551","href":"https://doi.org/10.82465/4551","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10294/15008","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Zilles, Sandra","Sadeqi, Mehdi"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Mouhoub, Malek"]},{"key":"dc:creator","label":"Author","values":["Alikhani, Majid"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-08-05T17:24:36Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-08-05T17:24:36Z"]},{"key":"dc:date.issued","label":"Date","values":["2021-09"]},{"key":"dc:publisher","label":"Institution","values":["Faculty of Graduate Studies and Research, University of Regina"]},{"key":"dc:type","label":"Dc Type","values":["master thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Master&apos;s"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MSc)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Faculty of Graduate Studies and Research, University of Regina"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.82465/4551"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10294/15008"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["A Thesis Submitted to the Faculty of Graduate Studies and Research In Partial Fulfillment of the Requirements for the Degree of Master of Science in Computer Science, University of Regina. ix, 149 p."]},{"key":"dc:description.abstract","label":"Abstract","values":["The field of Natural Language Processing (NLP) has witnessed substantial advancements due to both the development of new algorithms and the increase of computational capacities. As a result, novel NLP techniques have been successfully applied in various fields such as medicine, law, and finance. Companies and corporations publish several financial reports and disclosures every year to meet the requirements and needs of law enforcement entities, analysts, customers, and stakeholders. Processing these materials requires a tremendous amount of time and manual labor, especially with the unwieldy volume of data accumulated throughout the years. This thesis attempts to exploit the potential of modern NLP techniques to summarize financial reports in an automated fashion. We utilize an attention-based language model, namely RoBERTa to build a financially-oriented language model. This model is combined with various non-parametric methods to generate sentence vectors which can facilitate the efficient selection of important sentences in the document as part of the summarization algorithm. We illustrate the superiority of a financially tuned language model. Further, it is shown that simple non-parametric sentence embeddings achieve ROUGE scores that are comparable to or even better than those produced by Sentence Transformer and Universal Embeddings."]},{"key":"dc:title","label":"Title","values":["Automatic Summarization of Financial Reports"]}]}],"canonical_facts":{"dc:contributor.advisor":["Zilles, Sandra","Sadeqi, Mehdi"],"dc:contributor.committeemember":["Mouhoub, Malek"],"dc:creator":["Alikhani, Majid"],"dc:date.accessioned":["2022-08-05T17:24:36Z"],"dc:date.available":["2022-08-05T17:24:36Z"],"dc:date.issued":["2021-09"],"dc:description":["A Thesis Submitted to the Faculty of Graduate Studies and Research In Partial Fulfillment of the Requirements for the Degree of Master of Science in Computer Science, University of Regina. ix, 149 p."],"dc:description.abstract":["The field of Natural Language Processing (NLP) has witnessed substantial advancements due to both the development of new algorithms and the increase of computational capacities. As a result, novel NLP techniques have been successfully applied in various fields such as medicine, law, and finance. Companies and corporations publish several financial reports and disclosures every year to meet the requirements and needs of law enforcement entities, analysts, customers, and stakeholders. Processing these materials requires a tremendous amount of time and manual labor, especially with the unwieldy volume of data accumulated throughout the years. This thesis attempts to exploit the potential of modern NLP techniques to summarize financial reports in an automated fashion. We utilize an attention-based language model, namely RoBERTa to build a financially-oriented language model. This model is combined with various non-parametric methods to generate sentence vectors which can facilitate the efficient selection of important sentences in the document as part of the summarization algorithm. We illustrate the superiority of a financially tuned language model. Further, it is shown that simple non-parametric sentence embeddings achieve ROUGE scores that are comparable to or even better than those produced by Sentence Transformer and Universal Embeddings."],"dc:identifier.doi":["https://doi.org/10.82465/4551"],"dc:identifier.uri":["https://hdl.handle.net/10294/15008"],"dc:language.iso":["en"],"dc:publisher":["Faculty of Graduate Studies and Research, University of Regina"],"dc:title":["Automatic Summarization of Financial Reports"],"dc:type":["master thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Master&apos;s"],"thesis:degree_name":["Master of Science (MSc)"],"thesis:institution_name":["Faculty of Graduate Studies and Research, University of Regina"]},"updated_at":"2026-07-24T04:03:39Z"}