Diagnostic accuracy of the Ottawa 3DY and Short Blessed Test to detect cognitive dysfunction in geriatric patients presenting to the emergency department
Bibliographic record
Abstract
OBJECTIVES: Cognitive dysfunction (CD) is a common finding in geriatric patients presenting to the emergency department (ED). Our primary objective was to determine the diagnostic accuracy of the Ottawa 3DY (O3DY) and Short Blessed Test (SBT) as screening tools for the detection of CD in the ED. Our secondary objective was to estimate the inter-rater reliability of these instruments. METHODS: We conducted a prospective cross-sectional comparative study at an inner-city academic medical centre (annual ED visit census 86 000). Patients aged 75 years or greater were evaluated for inclusion, 163 were screened, 150 were deemed eligible and 117 were enrolled. The research team completed the O3DY, SBT and Mini-Mental State Exam (MMSE) for each participant. Descriptive statistics were calculated. Sensitivity and specificity of the O3DY and SBT were calculated in STATA V.11.2 using the MMSE as our criterion standard. RESULTS: We enrolled 117 patients from June to November 2016. The median ED length of stay at the time of completion of all tests was 1:40 (IQR 1:34-1:46). The sensitivity of the O3DY was 71.4% (95% CI 47.8 to 95.1), and specificity was 56.3% (46.7-65.9). Sensitivity of the SBT was 85.7% (67.4-99.9) and specificity was 58.3% (48.7-67.8). The receiver operating characteristic area under the curve was calculated for the O3DY (0.51; 95% CI 0.42 to 0.61) and SBT (0.52; 95% CI 0.43 to 0.61) relative to the MMSE. Inter-rater reliability for the O3DY (k=0.64) and SBT (k=0.63) were good. CONCLUSION: In a cohort of geriatric patients presenting to an inner-city academic ED, the O3DY and SBT tools demonstrate moderate sensitivity and specificity for the detection of CD. Inter-rater reliability for the O3DY and SBT were good. Future research on this topic should attempt to derive and validate ED-specific screening tools, which will hopefully result in more robust likelihood ratios for the screening of CD in ED geriatric patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.075 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".