The Graduate School and University Center of The City University of New York
Expanding the Corpus of Vocalized Hebrew Text: Compiling an Unvocalized Text Corpus and Building an Online Interface for Vocalization Annotation
Abstract
dc:description.abstract<p>Written modern Hebrew presents a unique challenge for training computational models for language processing because modern Hebrew text often lacks vocalization. The lack of available vocalized Hebrew data can lead to ambiguity in training these models and generally hinders work on natural language processing problems. The goal of this project is to contribute to the collection of vocalized Hebrew text by collecting and preprocessing a large corpus of unvocalized Hebrew text and building an online annotation tool. The annotation tool allows people to upload unvocalized Hebrew text, to annotate by adding Hebrew vocalization, and to download comma-separated values files of vocalized text. This project seeks to enhance the space both by collecting vocalized Hebrew text and by building the vocalization annotation interface—allowing for the continuous growth of a vocalized Hebrew text corpus.</p>
Degree
thesis:*- Name thesis:degree_name
- Master of Arts
- Level thesis:degree_level
- Master
- Discipline thesis:degree_discipline
- Linguistics
- Grantor
- The Graduate School and University Center of The City University of New York
- Year dc:date.available
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Bloch, Rachel Shanblatt
- Advisor dc:contributor.advisor
-
- Kyle Gorman
Subjects
dc:subject × 9Identifiers
dc:identifier.*- Repository record dc:identifier
- https://academicworks.cuny.edu/gc_etds/5891
- OAI identifier oai:identifier
- oai:academicworks.cuny.edu:gc_etds-6975