Validating a single-question depression measure among older adults
Bibliographic record
Abstract
BACKGROUND: A single-item depression measure may not be adequate in capturing the complex entity of mental health, despite wide use of this indicator in community studies. This study evaluated the accuracy of a single-question depression measure in comparison to two composite indices-the Center for Epidemiologic Studies Depression Scale (CESD) and the Geriatric Depression Scale (GDS). MATERIALS AND METHODS: A total of 800 elderly participants ranging from 60 to 89 years of age and residing in Seoul were recruited using a multistage sampling scheme in 2015. The survey was conducted by trained interviewers with a constructed questionnaire. Reliability and validity measures such as the Kappa index, sensitivity, specificity, PPV, NPV, and AUC were used to evaluate the accuracy of the single question measure. Socio-demographic group differences in accuracy were compared by age, sex, marital status, education, employment, and financial status. RESULTS: The prevalence of depression by a single-question measure was much lower than those of CESD and GDS (5.5%, 12.3%, and 12.1%, respectively). The sensitivity of the single-item measure, based on CESD and GDS, was extremely low at 30.6% and 36.1%. In the subgroup analysis, however, there was a marked educational discrepancy in all accuracy measures; in sensitivity, people with a university degree or higher showed about 2.4 times higher sensitivity than those having only a primary school education. CONCLUSIONS: The results show that a single-question depression measure should be used with caution. In addition, the single-question measure could substantially underestimate depression among the risk group of older adults.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".