{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/92859"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/92859","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Automatic generation of tunable analogy benchmarks for word representations","abstract":"We present a method to automatically generate syntactic analogy datasets for the evaluation of word representations in an unsupervised manner. The automatic generation also allows for customization in terms of word-frequencies, syntactic rules, part-of-speech tags and size of the dataset. We show the ability of our method to generate cross-lingual analogy task datasets for languages other than English, where evaluation datasets are limited if not nonexistent, by constructing datasets for French, German, Spanish, Arabic and Hebrew. Our method clusters pairs of words into morphological rules in an unsupervised manner, using which we generate analogy questions for different rules. We show the quality of an automatically generated dataset by checking the correlation of the performance of different word representations on it with the performance of the same representations on the Google analogy dataset. The values exhibited a high correlation of 95%. Moreover, we showcase the benefits of customization through studying the performance of different word representations when varying the frequency of words in the dataset.","abstract_html":"We present a method to automatically generate syntactic analogy datasets for the evaluation of word representations in an unsupervised manner. The automatic generation also allows for customization in terms of word-frequencies, syntactic rules, part-of-speech tags and size of the dataset. We show the ability of our method to generate cross-lingual analogy task datasets for languages other than English, where evaluation datasets are limited if not nonexistent, by constructing datasets for French, German, Spanish, Arabic and Hebrew. Our method clusters pairs of words into morphological rules in an unsupervised manner, using which we generate analogy questions for different rules. We show the quality of an automatically generated dataset by checking the correlation of the performance of different word representations on it with the performance of the same representations on the Google analogy dataset. The values exhibited a high correlation of 95%. Moreover, we showcase the benefits of customization through studying the performance of different word representations when varying the frequency of words in the dataset.","abstract_has_math":false,"creators":["Sakakini, Tarek J"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Viswanath, Pramod","Bhat, Suma"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-11-10T17:55:16Z","date_published":"2016-11-10T17:55:16Z","updated_at":"2026-07-22T22:26:35Z","subjects":["Natural Language Processing (NLP)","word representations","evaluation","semantics","morphology"],"languages":["en"],"rights":["Copyright 2016 Tarek Sakakini"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/92859","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Viswanath, Pramod","Bhat, Suma"]},{"key":"dc:creator","label":"Author","values":["Sakakini, Tarek J"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-11-10T17:55:16Z","2016-07-20","2016-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Processing (NLP)","word representations","evaluation","semantics","morphology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Tarek Sakakini"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/92859"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We present a method to automatically generate syntactic analogy datasets for the evaluation of word representations in an unsupervised manner. The automatic generation also allows for customization in terms of word-frequencies, syntactic rules, part-of-speech tags and size of the dataset. We show the ability of our method to generate cross-lingual analogy task datasets for languages other than English, where evaluation datasets are limited if not nonexistent, by constructing datasets for French, German, Spanish, Arabic and Hebrew. Our method clusters pairs of words into morphological rules in an unsupervised manner, using which we generate analogy questions for different rules. We show the quality of an automatically generated dataset by checking the correlation of the performance of different word representations on it with the performance of the same representations on the Google analogy dataset. The values exhibited a high correlation of 95%. Moreover, we showcase the benefits of customization through studying the performance of different word representations when varying the frequency of words in the dataset.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Tarek Sakakini, accepted the attached license on 2016-07-16 at 17:49.","The student, Tarek Sakakini, submitted this Thesis for approval on 2016-07-16 at 17:57.","This Thesis was approved for publication on 2016-07-20 at 08:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9976 on 2016-11-09 at 10:25:18","Made available in DSpace on 2016-11-10T17:55:16Z (GMT). No. of bitstreams: 2 SAKAKINI-THESIS-2016.pdf: 413983 bytes, checksum: 67cf2a00902f1ceeaedfbed6b19589f6 (MD5) LICENSE.txt: 4211 bytes, checksum: 5a6db85890b1b40467fded781bd85700 (MD5) Previous issue date: 2016-07-20"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Automatic generation of tunable analogy benchmarks for word representations"]}]}],"canonical_facts":{"dc:contributor":["Viswanath, Pramod","Bhat, Suma"],"dc:creator":["Sakakini, Tarek J"],"dc:date":["2016-11-10T17:55:16Z","2016-07-20","2016-08"],"dc:description":["We present a method to automatically generate syntactic analogy datasets for the evaluation of word representations in an unsupervised manner. The automatic generation also allows for customization in terms of word-frequencies, syntactic rules, part-of-speech tags and size of the dataset. We show the ability of our method to generate cross-lingual analogy task datasets for languages other than English, where evaluation datasets are limited if not nonexistent, by constructing datasets for French, German, Spanish, Arabic and Hebrew. Our method clusters pairs of words into morphological rules in an unsupervised manner, using which we generate analogy questions for different rules. We show the quality of an automatically generated dataset by checking the correlation of the performance of different word representations on it with the performance of the same representations on the Google analogy dataset. The values exhibited a high correlation of 95%. Moreover, we showcase the benefits of customization through studying the performance of different word representations when varying the frequency of words in the dataset.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Tarek Sakakini, accepted the attached license on 2016-07-16 at 17:49.","The student, Tarek Sakakini, submitted this Thesis for approval on 2016-07-16 at 17:57.","This Thesis was approved for publication on 2016-07-20 at 08:40.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9976 on 2016-11-09 at 10:25:18","Made available in DSpace on 2016-11-10T17:55:16Z (GMT). No. of bitstreams: 2 SAKAKINI-THESIS-2016.pdf: 413983 bytes, checksum: 67cf2a00902f1ceeaedfbed6b19589f6 (MD5) LICENSE.txt: 4211 bytes, checksum: 5a6db85890b1b40467fded781bd85700 (MD5) Previous issue date: 2016-07-20"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/92859"],"dc:language":["en"],"dc:rights":["Copyright 2016 Tarek Sakakini"],"dc:subject":["Natural Language Processing (NLP)","word representations","evaluation","semantics","morphology"],"dc:title":["Automatic generation of tunable analogy benchmarks for word representations"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:35Z"}