Effectiveness of Serious Games in Evaluating Cognitive Status of the Elderly: A Systematic Review and Meta-Analysis
Bibliographic record
Abstract
Early diagnosis of mild cognitive impairment (MCI) and Alzheimer's disease (AD) is very important in better management of these diseases, and serious games play an effective role in helping to diagnose these diseases more accurately owing to their innovative features. With respect to the diversity of available games, the purpose of this study was to investigate the effectiveness of using serious games to assess the cognitive status of the elderly at risk of MCI/AD. A systematic review was conducted and the correlation of serious game results with cognitive test scores were extracted from eligible studies for meta-analysis. We analyzed the correlation between the results of serious games with the scores of mini-mental state examination (MMSE), Addenbrooke's Cognitive Examination-revised edition (ACE-R), and Montreal Cognitive Assessment (MoCA) tests to evaluate cognitive status of the elderly at risk of MCI/AD, as well as the cognitive aspects examined by these tests. The random-effects model was used to obtain the overall correlation coefficient to assess the relationship between the results of serious games and the above mentioned paper-and-pencil tests. The correlation of game results with the MMSE, ACE-R, and MoCA was 0.604, 0.682, with 0.682, respectively. The correlation between the results of the games with the score of each cognitive aspect was also calculated. Overall, there is a positive correlation between serious game scores in terms of accurate patients' reactions with the scores of MMSE, ACE-R, and MoCA tests. Among the cognitive aspects, the highest correlation was obtained for fluency (0.591). For abstraction, however, the correlation was the lowest (0.036). In all three tests, the correlation was >0.6 and in cognitive aspects was <0.6. Thus, more studies should be conducted to develop serious games that are more in line with cognitive tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.012 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".