University of Illinois at Urbana-Champaign
Automatic generation of tunable analogy benchmarks for word representations
Abstract
dc:descriptionWe present a method to automatically generate syntactic analogy datasets for the evaluation of word representations in an unsupervised manner. The automatic generation also allows for customization in terms of word-frequencies, syntactic rules, part-of-speech tags and size of the dataset. We show the ability of our method to generate cross-lingual analogy task datasets for languages other than English, where evaluation datasets are limited if not nonexistent, by constructing datasets for French, German, Spanish, Arabic and Hebrew. Our method clusters pairs of words into morphological rules in an unsupervised manner, using which we generate analogy questions for different rules. We show the quality of an automatically generated dataset by checking the correlation of the performance of different word representations on it with the performance of the same representations on the Google analogy dataset. The values exhibited a high correlation of 95%. Moreover, we showcase the benefits of customization through studying the performance of different word representations when varying the frequency of words in the dataset.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2016
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Sakakini, Tarek J
- Contributors dc:contributor
-
- Viswanath, Pramod
- Bhat, Suma
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Copyright 2016 Tarek Sakakini
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/92859
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/92859