Back to results

National University of Singapore

Combining Speech with textual methods for arabic diacritization

Abstract

dc:description.abstract

The majority of studies on Arabic diacritization have employed textually inferred features alone. This thesis proposes a novel approach, where the weighted combination of speech with a text-based model is used to allow linguistically-insensitive acoustic information to correct and complement the errors generated by the text model's diacritic predictions. The acoustic model is based on Hidden Markov Models and the textual model on Conditional Random Fields. The combination brings significant reduction in error rates across all metrics, especially in case endings, which are the most difficult to predict. It gives results superior to those of conventional methods, with diacritic and word error rates of 1.6 and 5.2 inclusive of case endings, and 1.0 and 3.0 exclusive of them. Additionally, an interesting comparison is made between the diacritized solutions provided by two of the most popular morphological tools in the field of Arabic NLP, in the context of our combined system.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • AISHA SIDDIQA AZIM

Subjects

dc:subject × 1

Chain of custody

source
Harvested from
National University of Singapore
Base URL
scholarbank.nus.edu.sg/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

AISHA SIDDIQA AZIM. Combining Speech with textual methods for arabic diacritization. 2012.