DISSOCIATING RETEST EFFECTS FROM DEVELOPMENTAL CHANGE FOR PREDICTING COGNITIVE STATUS
Bibliographic record
Abstract
Abstract In longitudinal designs, unadjusted retest effects can confound developmental change estimates. This study utilized a measurement burst design and three-level multilevel modeling to a) independently parameterize short-term retest and long-term developmental change and b) employ these estimates as predictors of cognitive status at long-term follow-ups. Using data from Project MIND , participants (N=304; aged 64-92 years) were assessed across biweekly sessions nested within annual bursts (spanning up to 17 total assessments over four years). Cognitive impairment no dementia (CIND) status was classified at Years 4 (the final burst assessment) and 8 (the study end date). Response time inconsistencies (RTI) were computed to index intraindividual variability across RT trials of a one-back response time (BRT) task. Three-level multilevel models simultaneously yet independently estimated BRT RTI change across weeks and years, indexing short-term retest and long-term developmental change, respectively. Individual slope estimates were extracted and utilized in multinomial logistic regression models contrasting short- vs. long-term RTI change as predictors of long-term cognitive status. Results from the three-level models indicated that retest and developmental slopes yielded non-redundant sources of variance, providing unique estimates of change that would otherwise be confounded. Further, short- and long-term RTI differentially predicted cognitive status at Years 4 and 8; failing to benefit from retest effects on the BRT task was associated with increased likelihood of cognitive impairment. This innovative approach to parameterizing retest effects can reduce systematic bias in estimates of long-term developmental change, as well as highlight the utility of retest effects as predictors of cognitive health.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".