{"id":{"repo_id":"helsinki","oai_identifier":"oai:helda.helsinki.fi:10138/313475"},"canonical_url":"https://search.dev.ndltd.org/etd/helsinki/oai:helda.helsinki.fi:10138/313475","repository":{"repo_id":"helsinki","name":"University of Helsinki","base_url":"https://helda.helsinki.fi/server/oai/request"},"display":{"title":"Assessing text readability and quality with language models","abstract":"Automatic readability assessment is considered as a challenging task in NLP due to its high degree of subjectivity. The majority prior work in assessing readability has focused on identifying the level of education necessary for comprehension without the consideration of text quality, i.e., how naturally the text flows from the perspective of a native speaker. Therefore, in this thesis, we aim to use language models, trained on well-written prose, to measure not only text readability in terms of comprehension but text quality. In this thesis, we developed two word-level metrics based on the concordance of article text with predictions made using language models to assess text readability and quality. We evaluate both metrics on a set of corpora used for readability assessment or automated essay scoring (AES) by measuring the correlation between scores assigned by our metrics and human raters. According to the experimental results, our metrics are strongly correlated with text quality, which achieve 0.4-0.6 correlations on 7 out of 9 datasets. We demonstrate that GPT-2 surpasses other language models, including the bigram model, LSTM, and bidirectional LSTM, on the task of estimating text quality in a zero-shot setting, and GPT-2 perplexity-based measure is a reasonable indicator for text quality evaluation.","abstract_html":"Automatic readability assessment is considered as a challenging task in NLP due to its high degree of subjectivity. The majority prior work in assessing readability has focused on identifying the level of education necessary for comprehension without the consideration of text quality, i.e., how naturally the text flows from the perspective of a native speaker. Therefore, in this thesis, we aim to use language models, trained on well-written prose, to measure not only text readability in terms of comprehension but text quality. In this thesis, we developed two word-level metrics based on the concordance of article text with predictions made using language models to assess text readability and quality. We evaluate both metrics on a set of corpora used for readability assessment or automated essay scoring (AES) by measuring the correlation between scores assigned by our metrics and human raters. According to the experimental results, our metrics are strongly correlated with text quality, which achieve 0.4-0.6 correlations on 7 out of 9 datasets. We demonstrate that GPT-2 surpasses other language models, including the bigram model, LSTM, and bidirectional LSTM, on the task of estimating text quality in a zero-shot setting, and GPT-2 perplexity-based measure is a reasonable indicator for text quality evaluation.","abstract_has_math":false,"creators":["Liu, Yang"],"institution":"Helsingin yliopisto","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Helsingin yliopisto, Matemaattis-luonnontieteellinen tiedekunta","University of Helsinki, Faculty of Science","Helsingfors universitet, Matematisk-naturvetenskapliga fakulteten"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020","date_published":"2020","updated_at":"2026-07-27T19:56:27Z","subjects":[],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["URN:NBN:fi:hulib-202003191584"],"render_values":[{"text":"URN:NBN:fi:hulib-202003191584","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/10138/313475","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Helsingin yliopisto, Matemaattis-luonnontieteellinen tiedekunta","University of Helsinki, Faculty of Science","Helsingfors universitet, Matematisk-naturvetenskapliga fakulteten"]},{"key":"dc:creator","label":"Author","values":["Liu, Yang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2020"]},{"key":"dc:publisher","label":"Institution","values":["Helsingin yliopisto","University of Helsinki","Helsingfors universitet"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["URN:NBN:fi:hulib-202003191584","http://hdl.handle.net/10138/313475"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Automatic readability assessment is considered as a challenging task in NLP due to its high degree of subjectivity. The majority prior work in assessing readability has focused on identifying the level of education necessary for comprehension without the consideration of text quality, i.e., how naturally the text flows from the perspective of a native speaker. Therefore, in this thesis, we aim to use language models, trained on well-written prose, to measure not only text readability in terms of comprehension but text quality. In this thesis, we developed two word-level metrics based on the concordance of article text with predictions made using language models to assess text readability and quality. We evaluate both metrics on a set of corpora used for readability assessment or automated essay scoring (AES) by measuring the correlation between scores assigned by our metrics and human raters. According to the experimental results, our metrics are strongly correlated with text quality, which achieve 0.4-0.6 correlations on 7 out of 9 datasets. We demonstrate that GPT-2 surpasses other language models, including the bigram model, LSTM, and bidirectional LSTM, on the task of estimating text quality in a zero-shot setting, and GPT-2 perplexity-based measure is a reasonable indicator for text quality evaluation."]},{"key":"dc:title","label":"Title","values":["Assessing text readability and quality with language models"]}]}],"canonical_facts":{"dc:contributor":["Helsingin yliopisto, Matemaattis-luonnontieteellinen tiedekunta","University of Helsinki, Faculty of Science","Helsingfors universitet, Matematisk-naturvetenskapliga fakulteten"],"dc:creator":["Liu, Yang"],"dc:date.issued":["2020"],"dc:description.abstract":["Automatic readability assessment is considered as a challenging task in NLP due to its high degree of subjectivity. The majority prior work in assessing readability has focused on identifying the level of education necessary for comprehension without the consideration of text quality, i.e., how naturally the text flows from the perspective of a native speaker. Therefore, in this thesis, we aim to use language models, trained on well-written prose, to measure not only text readability in terms of comprehension but text quality. In this thesis, we developed two word-level metrics based on the concordance of article text with predictions made using language models to assess text readability and quality. We evaluate both metrics on a set of corpora used for readability assessment or automated essay scoring (AES) by measuring the correlation between scores assigned by our metrics and human raters. According to the experimental results, our metrics are strongly correlated with text quality, which achieve 0.4-0.6 correlations on 7 out of 9 datasets. We demonstrate that GPT-2 surpasses other language models, including the bigram model, LSTM, and bidirectional LSTM, on the task of estimating text quality in a zero-shot setting, and GPT-2 perplexity-based measure is a reasonable indicator for text quality evaluation."],"dc:identifier.uri":["URN:NBN:fi:hulib-202003191584","http://hdl.handle.net/10138/313475"],"dc:language.iso":["eng"],"dc:publisher":["Helsingin yliopisto","University of Helsinki","Helsingfors universitet"],"dc:title":["Assessing text readability and quality with language models"]},"updated_at":"2026-07-27T19:56:27Z"}