{"id":{"repo_id":"cuny-grad","oai_identifier":"oai:academicworks.cuny.edu:gc_etds-6975"},"canonical_url":"https://search.dev.ndltd.org/etd/cuny-grad/oai:academicworks.cuny.edu:gc_etds-6975","repository":{"repo_id":"cuny-grad","name":"City University of New York - Graduate Center","base_url":"https://academicworks.cuny.edu/do/oai/"},"display":{"title":"Expanding the Corpus of Vocalized Hebrew Text: Compiling an Unvocalized Text Corpus and Building an Online Interface for Vocalization Annotation","abstract":"<p>Written modern Hebrew presents a unique challenge for training computational models for language processing because modern Hebrew text often lacks vocalization. The lack of available vocalized Hebrew data can lead to ambiguity in training these models and generally hinders work on natural language processing problems. The goal of this project is to contribute to the collection of vocalized Hebrew text by collecting and preprocessing a large corpus of unvocalized Hebrew text and building an online annotation tool. The annotation tool allows people to upload unvocalized Hebrew text, to annotate by adding Hebrew vocalization, and to download comma-separated values files of vocalized text. This project seeks to enhance the space both by collecting vocalized Hebrew text and by building the vocalization annotation interface—allowing for the continuous growth of a vocalized Hebrew text corpus.</p>","abstract_html":"&lt;p&gt;Written modern Hebrew presents a unique challenge for training computational models for language processing because modern Hebrew text often lacks vocalization. The lack of available vocalized Hebrew data can lead to ambiguity in training these models and generally hinders work on natural language processing problems. The goal of this project is to contribute to the collection of vocalized Hebrew text by collecting and preprocessing a large corpus of unvocalized Hebrew text and building an online annotation tool. The annotation tool allows people to upload unvocalized Hebrew text, to annotate by adding Hebrew vocalization, and to download comma-separated values files of vocalized text. This project seeks to enhance the space both by collecting vocalized Hebrew text and by building the vocalization annotation interface—allowing for the continuous growth of a vocalized Hebrew text corpus.&lt;/p&gt;","abstract_has_math":false,"creators":["Bloch, Rachel Shanblatt"],"institution":"The Graduate School and University Center of The City University of New York","degree_name":"Master of Arts","degree_level":"Master","degree_discipline":"Linguistics","degree_department":null,"school":null,"contributors":[],"advisors":["Kyle Gorman"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-06-01T07:00:00Z","date_published":"2024-06-01T07:00:00Z","updated_at":"2026-07-24T02:00:26Z","subjects":["Computational Linguistics","Jewish Studies","Linguistics","Hebrew","natural language processing","language annotation","diacritization","vocalization","niqud"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://academicworks.cuny.edu/gc_etds/5891","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Kyle Gorman"]},{"key":"dc:creator","label":"Author","values":["Bloch, Rachel Shanblatt"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2024-05-02T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Linguistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Master"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Arts"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The Graduate School and University Center of The City University of New York"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computational Linguistics","Jewish Studies","Linguistics","Hebrew","natural language processing","language annotation","diacritization","vocalization","niqud"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://academicworks.cuny.edu/gc_etds/5891"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Written modern Hebrew presents a unique challenge for training computational models for language processing because modern Hebrew text often lacks vocalization. The lack of available vocalized Hebrew data can lead to ambiguity in training these models and generally hinders work on natural language processing problems. The goal of this project is to contribute to the collection of vocalized Hebrew text by collecting and preprocessing a large corpus of unvocalized Hebrew text and building an online annotation tool. The annotation tool allows people to upload unvocalized Hebrew text, to annotate by adding Hebrew vocalization, and to download comma-separated values files of vocalized text. This project seeks to enhance the space both by collecting vocalized Hebrew text and by building the vocalization annotation interface—allowing for the continuous growth of a vocalized Hebrew text corpus.</p>"]},{"key":"dc:title","label":"Title","values":["Expanding the Corpus of Vocalized Hebrew Text: Compiling an Unvocalized Text Corpus and Building an Online Interface for Vocalization Annotation"]}]}],"canonical_facts":{"dc:contributor.advisor":["Kyle Gorman"],"dc:creator":["Bloch, Rachel Shanblatt"],"dc:date.available":["2024-05-02T07:00:00Z"],"dc:description.abstract":["<p>Written modern Hebrew presents a unique challenge for training computational models for language processing because modern Hebrew text often lacks vocalization. The lack of available vocalized Hebrew data can lead to ambiguity in training these models and generally hinders work on natural language processing problems. The goal of this project is to contribute to the collection of vocalized Hebrew text by collecting and preprocessing a large corpus of unvocalized Hebrew text and building an online annotation tool. The annotation tool allows people to upload unvocalized Hebrew text, to annotate by adding Hebrew vocalization, and to download comma-separated values files of vocalized text. This project seeks to enhance the space both by collecting vocalized Hebrew text and by building the vocalization annotation interface—allowing for the continuous growth of a vocalized Hebrew text corpus.</p>"],"dc:identifier":["https://academicworks.cuny.edu/gc_etds/5891"],"dc:subject":["Computational Linguistics","Jewish Studies","Linguistics","Hebrew","natural language processing","language annotation","diacritization","vocalization","niqud"],"dc:title":["Expanding the Corpus of Vocalized Hebrew Text: Compiling an Unvocalized Text Corpus and Building an Online Interface for Vocalization Annotation"],"thesis:degree_discipline":["Linguistics"],"thesis:degree_level":["Master"],"thesis:degree_name":["Master of Arts"],"thesis:institution_name":["The Graduate School and University Center of The City University of New York"]},"updated_at":"2026-07-24T02:00:26Z"}