{"id":{"repo_id":"nus","oai_identifier":"oai:scholarbank.nus.edu.sg:10635/238651"},"canonical_url":"https://search.dev.ndltd.org/etd/nus/oai:scholarbank.nus.edu.sg:10635/238651","repository":{"repo_id":"nus","name":"National University of Singapore","base_url":"https://scholarbank.nus.edu.sg/oai/request"},"display":{"title":"SELF-ENHANCED VOCABULARY LEARNING LATENT DIRECHLET ALLOCATION","abstract":"Many studies have been conducted to identify valuable information with sentiments and classifications. In the latter, topics models such as Latent Dirichlet Allocation (LDA) (Blei et al., 2003) and Biterm Topic Model (BTM) (Yan et al., 2013) are conventional probabilistic models designed to unveil latent topic structure within texts. In the paper, we proposed a novel topic model to process short texts called Self-Enhanced Vocabulary Learning Latent Dirichlet Allocation (SVL-LDA), which is an extension from LDA by incorporating rolling window training of parameters of prior distribution to handle the sparsity problem. Especially, a larger base of information is utilized to finetune the corpus-level parameters. Then the LDA module will identify the coherence of topics-documents and word-topics. Empirical results show that our approach appears to outperform the baseline methods under certain conditions and demonstrate the practical application to social media.","abstract_html":"Many studies have been conducted to identify valuable information with sentiments and classifications. In the latter, topics models such as Latent Dirichlet Allocation (LDA) (Blei et al., 2003) and Biterm Topic Model (BTM) (Yan et al., 2013) are conventional probabilistic models designed to unveil latent topic structure within texts. In the paper, we proposed a novel topic model to process short texts called Self-Enhanced Vocabulary Learning Latent Dirichlet Allocation (SVL-LDA), which is an extension from LDA by incorporating rolling window training of parameters of prior distribution to handle the sparsity problem. Especially, a larger base of information is utilized to finetune the corpus-level parameters. Then the LDA module will identify the coherence of topics-documents and word-topics. Empirical results show that our approach appears to outperform the baseline methods under certain conditions and demonstrate the practical application to social media.","abstract_has_math":false,"creators":["HUANG YI HSIANG"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08-31","date_published":"2022-08-31","updated_at":"2026-07-24T03:33:34Z","subjects":["tweets, LDA, BTM, NLP, topic, text"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["HUANG YI HSIANG"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2022-08-31"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://scholarbank.nus.edu.sg/handle/10635/238651"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["tweets, LDA, BTM, NLP, topic, text"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarbank.nus.edu.sg/bitstreams/cf022dc9-d34c-4802-b40b-f74ef7d61aa3/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Many studies have been conducted to identify valuable information with sentiments and classifications. In the latter, topics models such as Latent Dirichlet Allocation (LDA) (Blei et al., 2003) and Biterm Topic Model (BTM) (Yan et al., 2013) are conventional probabilistic models designed to unveil latent topic structure within texts. In the paper, we proposed a novel topic model to process short texts called Self-Enhanced Vocabulary Learning Latent Dirichlet Allocation (SVL-LDA), which is an extension from LDA by incorporating rolling window training of parameters of prior distribution to handle the sparsity problem. Especially, a larger base of information is utilized to finetune the corpus-level parameters. Then the LDA module will identify the coherence of topics-documents and word-topics. Empirical results show that our approach appears to outperform the baseline methods under certain conditions and demonstrate the practical application to social media."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["4f3788ad44339484a460098a1a6d4054","d9c28385e06fa086518acbd4f218a1b7"]},{"key":"dc:title","label":"Title","values":["SELF-ENHANCED VOCABULARY LEARNING LATENT DIRECHLET ALLOCATION"]}]}],"canonical_facts":{"dc:creator":["HUANG YI HSIANG"],"dc:date.issued":["2022-08-31"],"dc:description.abstract":["Many studies have been conducted to identify valuable information with sentiments and classifications. In the latter, topics models such as Latent Dirichlet Allocation (LDA) (Blei et al., 2003) and Biterm Topic Model (BTM) (Yan et al., 2013) are conventional probabilistic models designed to unveil latent topic structure within texts. In the paper, we proposed a novel topic model to process short texts called Self-Enhanced Vocabulary Learning Latent Dirichlet Allocation (SVL-LDA), which is an extension from LDA by incorporating rolling window training of parameters of prior distribution to handle the sparsity problem. Especially, a larger base of information is utilized to finetune the corpus-level parameters. Then the LDA module will identify the coherence of topics-documents and word-topics. Empirical results show that our approach appears to outperform the baseline methods under certain conditions and demonstrate the practical application to social media."],"dc:format.checksum.md5":["4f3788ad44339484a460098a1a6d4054","d9c28385e06fa086518acbd4f218a1b7"],"dc:identifier.uri":["https://scholarbank.nus.edu.sg/bitstreams/cf022dc9-d34c-4802-b40b-f74ef7d61aa3/download"],"dc:relation.isreferencedby":["https://scholarbank.nus.edu.sg/handle/10635/238651"],"dc:subject":["tweets, LDA, BTM, NLP, topic, text"],"dc:title":["SELF-ENHANCED VOCABULARY LEARNING LATENT DIRECHLET ALLOCATION"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T03:33:34Z"}