Mapping of validated apathy scales onto the apathy diagnostic criteria for neurocognitive disorders
Bibliographic record
Abstract
BACKGROUND: Diagnostic criteria for apathy in neurocognitive disorders (DCA-NCD) have recently been updated. OBJECTIVES: We investigated whether validated scales measuring apathy severity capture the three dimensions of the DCA-NCD (diminished initiative, diminished interest, diminished emotional expression). MEASUREMENTS: Degree of mapping ("not at all", "weakly", or "strongly") between items on two commonly used apathy scales, the Neuropsychiatric Inventory-Clinician (NPI-C) apathy and Apathy Evaluation Scale (AES), with the DCA-NCD overall and its 3 dimensions was evaluated by survey. DESIGN: Survey participants, either experts (n = 12, DCA-NCD authors) or scientific community members (n = 19), rated mapping for each item and mean scores were calculated. Interrater reliability between expert and scientific community members was assessed using Cohen's kappa. RESULTS: According to experts, 9 of 11 (81.8%) NPI-C apathy items and 6 of 18 (33.3%) AES items mapped strongly onto the DCA-NCD overall. For the scientific community group, 10 of 11 (90.9%) NPI-C apathy items and 7 of 18 (38.8%) AES items mapped strongly onto the DCA-NCD overall. The overall mean mapping scores were higher for the NPI-C apathy compared to the AES for both expert (t (11) = 3.13, p = .01) and scientific community (t (17) = 3.77, p = .002) groups. There was moderate agreement between the two groups on overall mapping for the NPI-C apathy (kappa= 0.74 (0.57, 1.00)) and AES (kappa= 0.63 (0.35, 1.00)). CONCLUSIONS: More NPI-C apathy than AES items mapped strongly and uniquely onto the DCA-NCD and its dimensions. The NPI-C apathy may better capture the DCA-NCD and its dimensions compared with the AES.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".