{"id":{"repo_id":"auckland-ms","oai_identifier":"oai:researchspace.auckland.ac.nz:2292/65940"},"canonical_url":"https://search.dev.ndltd.org/etd/auckland-ms/oai:researchspace.auckland.ac.nz:2292/65940","repository":{"repo_id":"auckland-ms","name":"University of Auckland","base_url":"https://researchspace.auckland.ac.nz/server/oai/request"},"display":{"title":"Improved Techniques for Identifying Software Product Development Insights","abstract":"Understanding the needs of users is essential for developing quality software products. This can be done using requirements engineering (RE), where software development teams consult product stakeholders regarding their requirements for the product. Such consultations can often be expensive, and so requirements engineers have looked to perform RE by using online user feedback, which can come in the form of app reviews, Tweets, and forum posts. Previous research has shown that this information can contain large amounts of requirements-relevant information that can be used in the development cycle (e.g. bug reports or feature requests). However, much of this data is unstructured and vast in scale. Therefore, recent research has explored using machine learning to automate some of the analysis of this data, such as using classifiers to automatically classify feedback into one of a set of categories (such as bug report or feature request), and clustering feedback into semantically coherent clusters that all mention the same bug or feature. While previous classifiers have proven accurate in some situations, they also exhibit low classification accuracy when applied to new data domains, limiting their applicability to requirements engineers. Similarly, previous clustering tools have presented groups of text that are similar based on characteristic irrelevant to software engineering. This thesis builds on previous work by demonstrating that both pre-training and multidataset fine-tuning can lead to higher user feedback accuracy. This thesis also finds ways of clustering text such that the clusters are grouped based on software engineering characteristics. We further test several methods of summarizing these clusters to find which is most able to communicate requirements relevant information to engineers. Our contributions combine to provide requirements engineers with a better understanding of how to create tools that can gain development insights from their user feedback, ultimately leading to higher quality software.","abstract_html":"Understanding the needs of users is essential for developing quality software products. This can be done using requirements engineering (RE), where software development teams consult product stakeholders regarding their requirements for the product. Such consultations can often be expensive, and so requirements engineers have looked to perform RE by using online user feedback, which can come in the form of app reviews, Tweets, and forum posts. Previous research has shown that this information can contain large amounts of requirements-relevant information that can be used in the development cycle (e.g. bug reports or feature requests). However, much of this data is unstructured and vast in scale. Therefore, recent research has explored using machine learning to automate some of the analysis of this data, such as using classifiers to automatically classify feedback into one of a set of categories (such as bug report or feature request), and clustering feedback into semantically coherent clusters that all mention the same bug or feature. While previous classifiers have proven accurate in some situations, they also exhibit low classification accuracy when applied to new data domains, limiting their applicability to requirements engineers. Similarly, previous clustering tools have presented groups of text that are similar based on characteristic irrelevant to software engineering. This thesis builds on previous work by demonstrating that both pre-training and multidataset fine-tuning can lead to higher user feedback accuracy. This thesis also finds ways of clustering text such that the clusters are grouped based on software engineering characteristics. We further test several methods of summarizing these clusters to find which is most able to communicate requirements relevant information to engineers. Our contributions combine to provide requirements engineers with a better understanding of how to create tools that can gain development insights from their user feedback, ultimately leading to higher quality software.","abstract_has_math":false,"creators":["Devine, Peter"],"institution":"ResearchSpace@Auckland","degree_name":"PhD","degree_level":"Doctoral","degree_discipline":"Software Engineering","degree_department":null,"school":null,"contributors":[],"advisors":["Blincoe, Kelly","Koh, Yun Sing"],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023","date_published":"2023","updated_at":"2026-07-24T01:06:10Z","subjects":[],"languages":[],"rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"rights_urls":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2292/65940","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Blincoe, Kelly","Koh, Yun Sing"]},{"key":"dc:creator","label":"Author","values":["Devine, Peter"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-09-20T02:26:33Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-09-20T02:26:33Z"]},{"key":"dc:date.issued","label":"Date","values":["2023"]},{"key":"dc:publisher","label":"Institution","values":["ResearchSpace@Auckland"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["UoA"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Software Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The University of Auckland"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/2292/65940"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Understanding the needs of users is essential for developing quality software products. This can be done using requirements engineering (RE), where software development teams consult product stakeholders regarding their requirements for the product. Such consultations can often be expensive, and so requirements engineers have looked to perform RE by using online user feedback, which can come in the form of app reviews, Tweets, and forum posts. Previous research has shown that this information can contain large amounts of requirements-relevant information that can be used in the development cycle (e.g. bug reports or feature requests). However, much of this data is unstructured and vast in scale. Therefore, recent research has explored using machine learning to automate some of the analysis of this data, such as using classifiers to automatically classify feedback into one of a set of categories (such as bug report or feature request), and clustering feedback into semantically coherent clusters that all mention the same bug or feature. While previous classifiers have proven accurate in some situations, they also exhibit low classification accuracy when applied to new data domains, limiting their applicability to requirements engineers. Similarly, previous clustering tools have presented groups of text that are similar based on characteristic irrelevant to software engineering. This thesis builds on previous work by demonstrating that both pre-training and multidataset fine-tuning can lead to higher user feedback accuracy. This thesis also finds ways of clustering text such that the clusters are grouped based on software engineering characteristics. We further test several methods of summarizing these clusters to find which is most able to communicate requirements relevant information to engineers. Our contributions combine to provide requirements engineers with a better understanding of how to create tools that can gain development insights from their user feedback, ultimately leading to higher quality software."]},{"key":"dc:title","label":"Title","values":["Improved Techniques for Identifying Software Product Development Insights"]}]}],"canonical_facts":{"dc:contributor.advisor":["Blincoe, Kelly","Koh, Yun Sing"],"dc:creator":["Devine, Peter"],"dc:date.accessioned":["2023-09-20T02:26:33Z"],"dc:date.available":["2023-09-20T02:26:33Z"],"dc:date.issued":["2023"],"dc:description.abstract":["Understanding the needs of users is essential for developing quality software products. This can be done using requirements engineering (RE), where software development teams consult product stakeholders regarding their requirements for the product. Such consultations can often be expensive, and so requirements engineers have looked to perform RE by using online user feedback, which can come in the form of app reviews, Tweets, and forum posts. Previous research has shown that this information can contain large amounts of requirements-relevant information that can be used in the development cycle (e.g. bug reports or feature requests). However, much of this data is unstructured and vast in scale. Therefore, recent research has explored using machine learning to automate some of the analysis of this data, such as using classifiers to automatically classify feedback into one of a set of categories (such as bug report or feature request), and clustering feedback into semantically coherent clusters that all mention the same bug or feature. While previous classifiers have proven accurate in some situations, they also exhibit low classification accuracy when applied to new data domains, limiting their applicability to requirements engineers. Similarly, previous clustering tools have presented groups of text that are similar based on characteristic irrelevant to software engineering. This thesis builds on previous work by demonstrating that both pre-training and multidataset fine-tuning can lead to higher user feedback accuracy. This thesis also finds ways of clustering text such that the clusters are grouped based on software engineering characteristics. We further test several methods of summarizing these clusters to find which is most able to communicate requirements relevant information to engineers. Our contributions combine to provide requirements engineers with a better understanding of how to create tools that can gain development insights from their user feedback, ultimately leading to higher quality software."],"dc:identifier.uri":["https://hdl.handle.net/2292/65940"],"dc:publisher":["ResearchSpace@Auckland"],"dc:relation.isreferencedby":["UoA"],"dc:rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"dc:rights.uri":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"dc:title":["Improved Techniques for Identifying Software Product Development Insights"],"dc:type":["Thesis"],"thesis:degree_discipline":["Software Engineering"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["PhD"],"thesis:institution_name":["The University of Auckland"]},"updated_at":"2026-07-24T01:06:10Z"}