PHQ-8 scores and estimation of depression prevalence
Bibliographic record
Abstract
In the Article by Jorge Arias-de la Torre and colleagues,1Arias-de la Torre J Vilagut G Ronaldson A et al.Prevalence and variability of current depressive disorder in 27 European countries: a population-based study.Lancet Public Health. 2021; 6: e729-e738Summary Full Text Full Text PDF PubMed Scopus (17) Google Scholar the authors used data for 258 888 individuals obtained from the second wave of the European Health Interview Survey to estimate depression prevalence in 27 European countries based on scores of 10 or higher on the eight-item Patient Health Questionnaire (PHQ-8). The authors reported an overall prevalence of current depressive disorder of 6·38% (95% CI 6·24–6·52) with substantial heterogeneity across countries.1Arias-de la Torre J Vilagut G Ronaldson A et al.Prevalence and variability of current depressive disorder in 27 European countries: a population-based study.Lancet Public Health. 2021; 6: e729-e738Summary Full Text Full Text PDF PubMed Scopus (17) Google Scholar Depression symptom questionnaires and standard cutoffs, such as PHQ-8 scores of 10 and higher, are not intended to estimate disorder prevalence, but are designed for screening purposes; they are intended to identify a higher number of individuals than would be diagnosed with depression if assessed using validated diagnostic criteria.2Thombs BD Kwakkenbos L Levis AW Benedetti A Addressing overestimating of the prevalence of depression based on self-report screening questionnaires.CMAJ. 2018; 190: e44-e49Crossref PubMed Scopus (69) Google Scholar An individual participant data meta-analysis of 44 studies,3Levis B Benedetti A Ioannidis JPA et al.Patient Health Questionnaire-9 scores do not accurately estimate depression prevalence: individual participant data meta-analysis.J Clin Epidemiol. 2020; 122: 115-128Summary Full Text Full Text PDF Scopus (54) Google Scholar which included 9242 participants (of whom 1389 had Structured Clinical Interview for DSM [SCID] major depression) found that, on average, prevalence based on nine-item Patient Health Questionnaire (PHQ-9) scores of 10 or higher (which perform similarly to scores of 10 or higher on the PHQ-84Wu Y Levis B Riehm KE et al.Equivalency of the diagnostic accuracy of the PHQ-8 and PHQ-9: a systematic review and individual participant data meta-analysis.Psychol Med. 2020; 50: 1368-1380Crossref PubMed Scopus (47) Google Scholar) overestimated SCID-based prevalence by 11·9%. In the 44 studies, the mean ratio of PHQ-9 scores of 10 or higher to SCID-based prevalence was 2·5.3Levis B Benedetti A Ioannidis JPA et al.Patient Health Questionnaire-9 scores do not accurately estimate depression prevalence: individual participant data meta-analysis.J Clin Epidemiol. 2020; 122: 115-128Summary Full Text Full Text PDF Scopus (54) Google Scholar Consistent with evidence that PHQ-9 scores of 10 or higher exaggerate prevalence, although the PHQ-8 assesses symptoms in the previous 2 weeks, in the European Health Interview Survey, the prevalence of current depressive disorder was higher than 12-month European prevalence based on a validated diagnostic interview.5Alonso J Angermeyer MC Bernert S et al.Prevalence of mental disorders in Europe: results from the European Study of the Epidemiology of Mental Disorders (ESEMeD) project.Acta Psychiatr Scand. 2004; 109: 21-27Google Scholar Depression is an important concern. However, using the proportion of individuals with scores above screening cutoffs on self-report questionnaires does not generate valid prevalence estimates, and the estimates reported by Arias-de la Torre and colleagues are not likely to represent the actual prevalence of depression in Europe. There are ways to incorporate self-report questionnaires into methods for estimating prevalence, but simply reporting the proportion of participants with positive screens is not recommended.2Thombs BD Kwakkenbos L Levis AW Benedetti A Addressing overestimating of the prevalence of depression based on self-report screening questionnaires.CMAJ. 2018; 190: e44-e49Crossref PubMed Scopus (69) Google Scholar We declare no competing interests. Prevalence and variability of current depressive disorder in 27 European countries: a population-based studyDepressive disorders, although common across Europe, vary substantially in prevalence between countries. These results could be a baseline for monitoring the prevalence of current depressive disorder both at a country level in Europe and for planning health-care resources and services. Full-Text PDF Open AccessPHQ-8 scores and estimation of depression prevalence – Author's replyWe thank Brooke Levis and colleagues for their interest in our work and for suggesting that we might have overestimated the prevalence of depression by using the eight-item Patient Health Questionnaire (PHQ-8) in our study.1 Although we acknowledged the limitations associated with the use of the PHQ-8, we believe that further discussion is required. Full-Text PDF Open Access
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".