Simplifying Complex Figure scoring: Data from the Emory Healthy Brain Study and initial clinical validation
Bibliographic record
Abstract
Abstract Objective: To introduce the Emory 10-element Complex Figure (CF) scoring system and recognition task. We evaluated the relationship between Emory CF scoring and traditional Osterrieth CF scoring approach in cognitively healthy volunteers. Additionally, a cohort of patients undergoing deep brain stimulation (DBS) evaluation was assessed to compare the scoring methods in a clinical population. Method: The study included 315 volunteers from the Emory Healthy Brain Study (EHBS) with Montreal Cognitive Assessment (MoCA) scores of 24/30 or higher. The clinical group consisted of 84 DBS candidates. Scoring time differences were analyzed in a subset of 48 DBS candidates. Results: High correlations between scoring methods were present for non-recognition components in both cohorts (EHBS: Copy r = 0.76, Immediate r = 0.86, Delayed r = 0.85, Recognition r = 47; DBS: Copy r = 0.80, Immediate r = 0.84, Delayed Recall r = 0.85, Recognition r = 0.37). Emory CF scoring times were significantly shorter than Osterrieth times across non-recognition conditions (all p < 0.00001, individual Cohen’s d: 1.4–2.4), resulting in an average time savings of 57%. DBS patients scored lower than EHBS participants across CF memory measures, with larger effect sizes for Emory CF scoring (Cohen’s d range = 1.0–1.2). Emory CF scoring demonstrated better group classification in logistic regression models, improving DBS candidate classification from 16.7% to 32.1% compared to Osterrieth scoring. Conclusions: Emory CF scoring yields results that are highly correlated with traditional Osterrieth scoring, significantly reduces scoring time burden, and demonstrates greater sensitivity to memory decline in DBS candidates. Its efficiency and sensitivity make Emory CF scoring well-suited for broader implementation in clinical research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".