Digital Ergonomics of NavegApp, a Novel Serious Game for Spatial Cognition Assessment: Content Validity and Usability Study
Bibliographic record
Abstract
Background Alzheimer disease (AD) is the leading cause of dementia worldwide. With aging populations and limited access to effective treatments, there is an urgent need for innovative markers to support timely preventive interventions. Emerging evidence highlights spatial cognition (SC) as a valuable source of cognitive markers for AD. This study presents NavegApp, a serious game (SG) designed to assess 3 key components of SC, which show potential as cognitive markers for the early detection of AD. Objective This study aimed to determine the content validity and usability perception of NavegApp across multiple groups of interest. Methods A multistep process integrating methodologies from software engineering, psychometrics, and health measurement was implemented to validate the software. Our approach was structured into 3 stages, guided by the software life cycle for health and the Consensus-Based Standards for the Selection of Health Status Measurement Instruments (COSMIN) recommendations for evaluating the psychometric quality of health instruments. To assess content validity, a panel of 8 experts evaluated the relevance and representativeness of tasks included in the app. In addition, 212 participants, categorized into 5 groups based on their clinical status and risk level for AD, were recruited to evaluate the app’s digital ergonomics and usability at various stages of development. Complementary analyses were performed to identify group differences and to explore the association between task difficulty and user agreeableness. Results NavegApp was validated as a highly usable tool by both experts and users. The expert panel confirmed that the tasks included in the game were representative (Aiken V=0.96-1.00) and relevant (Aiken V=0.96-1.00) for measuring SC components. Both experts and nonexperts rated NavegApp’s digital ergonomics positively, with minimal differences between groups (rrb 0.08-0.29). Differences in usability perceptions were observed among participants with sporadic mild cognitive impairment compared to cognitively healthy individuals (rrb 0.26-0.29). A moderate association was also identified between task difficulty and user agreeableness (Cramér V=0.37, 95% CI 0.28-0.54). Conclusions NavegApp is a valid and user-friendly SG designed for SC assessment, developed by integrating software engineering and psychometric evaluation methodologies. While the results are promising, further studies are warranted to evaluate its diagnostic accuracy and construct validity. This work outlines a comprehensive framework for SG development in cognitive assessment, emphasizing the importance of incorporating psychometric validity measures from the outset of the design process.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".