An interplay between cross-cultural and psychometric factors in the Montreal Cognitive Assessment: Experience from the language of a small nation
Bibliographic record
Abstract
The present study aimed to test the hypothesis that the total word length on the Memory subtest of the Czech version of the MoCA, which is 12 syllables compared to the English version of 7 syllables, would have a significant effect on Delayed Recall scores compared to the newly created well-balanced version of the test (further MoCA-WLE). In the original Czech version of MoCA, we replaced the 12-syllable word list in the Memory subtest with a 7-syllable list (MoCA-WLE) to make it equivalent to the standard English version in this respect. We analyzed data from 83 participants in the original MoCA group (70.63 ± 7.01 years old, 14.61 ± 3.17 years of education, 30.12% males) and 83 participants in the MoCA-WLE group (70.72 ± 6.95 years old, 14.93 ± 3.48 years of education, 30.12% males). We did not find evidence for a significant word-length effect in the original MoCA versus MoCA-WLE Delayed Recall in either the Mann–Whitney U test (W = 3418.0, p = .932) or multilevel binomial regression (b = 0.10, 95% posterior probability interval [−0.46, 0.68]). The present study shows cross-cultural limits in the adaptation of the test material. The results underline the caveats of such an approach to test adaptation. Fortunately, 12-syllables in the MoCA Memory Czech version versus the original 7-syllable list did not show a detectable word-length effect. We did not find evidence for differential item functioning or cultural item bias. The original MoCA Czech version is psychometrically comparable to the original English version.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".