Bibliographic record
Abstract
Tigrinya Analogy Test for evaluating Word Embeddings This is a Tigrinya version of the Google Analogy Test set, which is used to evaluate English word-embedding models. The analogy test is a well-established strategy to empirically evaluate the quality of word-embedding models. More information about the English task can be found at the ACL Wiki. This data is was first machine translated then manually verified by a native speaker to reduce errors. Some aspects of the original analogy test is focused on English and may not transfer well to other languages, such as those related to grammar or morphology. Therefore, we have discarded examples that became irrelevant in Tigrinya when adapting the task. Finally, there are a total of 18465 entries in the Tigrinya Analogy Test set, while the source English data has 19544 entries. An entry is dropped if the translations led to one of the following conditions: If the source word pair map to one Tigrinya word, for example, lucky & luckiest both correspond to ዕድለኛ. If the source word results in a multi-word expression. For example, grandson (ወዲ ጓል / ወዲ ወዲ), granddaughter (ጓል ጓል / ጓል ወዲ). This because the typical word-embedding approaches such as word2vec are not designed to predict multi-word phrases. Test Sections The test includes a series of semantic and syntactic analogies divided up into subsections including world capitals, currencies, family, tense, and plurality. The test contains the following sections: capital-world currency city-in-state family gram1-adjective-to-adverb gram2-opposite gram3-comparative gram4-superlative gram5-present-participle gram6-nationality-adjective gram7-past-tense gram8-plural gram9-plural-verbs Examples: Semantic section of World Capitals: “ኣስመራ: ኤርትራ as ፓሪስ: ?” and if the model responds correctly it will return: “ፈረንሳ”. Semantic section of Family section: “ሰብኣይ: ሰበይቲ as ወዲ: ጓል”. Syntax section with tense, a sample analogy might be “Walk: Walked as Run: Ran”. Evaluation The final accuracy of a model is the proportion of the questions that the model answers correctly. Generally, a better-quality model would answer more questions correctly than a model of lower quality. However, note that a model with low performance on this analogy test, might still contain useful information, but may not be robust or good enough for more complex tasks. Limitations The analogy test could be a good indicator of the quality of word-embeddings, but it should be used with caution when comparing models trained on varying domains of data. It shall not be expected to generalize equally to all domains. The final score can be affected by the size, vocabulary, and domain of the text with which the models are trained on. For example, this may not be a good benchmark to compare models trained on news text vs posts on social media. Even though a manual sanity check was performed, we note that the semi-automatic construction of the Tigrinya test set might contains errors. If you discover any, you are welcome to contribute back by either opening an Issue at the GitHub repo, https://github.com/fgaim/tigrinya-analogy-test. Citation If you use this resource in your research, please cite it accordingly.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.041 |
| Meta-epidemiology (narrow) | 0.003 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.020 | 0.017 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".