Predictive validity of the Oxford digital multiple errands test (OxMET) for functional outcomes after stroke
Bibliographic record
Abstract
The Oxford Digital Multiple Errands Test (OxMET) is a brief computer-tablet based cognitive screen, intended as an ecologically valid assessment of executive dysfunction. We examined aspects of predictive validity in relation to functional outcomes. Participants (≤ 2 months post-stroke) were recruited from an English-speaking stroke rehabilitation in-patient setting. Participants completed OxMET. The Barthel Index, Therapy Outcome Measure (TOMS), and modified Rankin Scale (mRS) were collected from medical notes. Participants were followed up after 6-months and completed the Nottingham Extended Activities of Daily Living (NEADL) scale. 117 participants were recruited (M = 26.18 days post-stroke (SD = 25.16), mean 74.44yrs (SD = 12.88), median NIHSS 8.32 (IQR = 5-11)). Sixty-six completed a follow-up (M = 73.94yrs (SD = 12.68), median NIHSS 8 (IQR = 4-11)). Significant associations were found between TOMS and mRS. At 6-month follow up, we found a moderate predictive relationship between the OxMET accuracy and NEADL (R2 = .29, p < .001), and we did not find this prediction with MoCA taken at 6-months. The subacute OxMET associated with measures of functionality and disability in a rehabilitation context, and in activities of daily living. The OxMET is an assessment of executive function with good predictive validity on clinically relevant functional outcome measures that may be more predictive than other cognitive tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.019 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".