Diagnostic accuracy of case-finding questions to identify perinatal depression
Bibliographic record
Abstract
BACKGROUND: Guidelines for perinatal mental health care recommend the use of two case-finding questions about depressed feelings and loss of interest in activities, despite the absence of validation studies in this context. We examined the diagnostic accuracy of these questions and of a third question about the need for help asked of women receiving perinatal care. METHODS: We evaluated self-reported responses to two case-finding questions against an interviewer-assessed diagnostic standard (DSM-IV criteria for major depressive disorder) among 152 women receiving antenatal care at 26-28 weeks' gestation and postnatal care at 5-13 weeks after delivery. Among women who answered "yes" to either question, we assessed the usefulness of asking a third question about the need for help. We calculated sensitivity, specificity and likelihood ratios for the two case-finding questions and for the added question about the need for help. RESULTS: Antenatally, the two case-finding questions had a sensitivity of 100% (95% confidence interval [CI] 77%-100%), a specificity of 68% (95% CI 58%-76%), a positive likelihood ratio of 3.03 (95% CI 2.28-4.02) and a negative likelihood ratio of 0.041 (95% CI 0.003-0.63) in identifying perinatal depression. Postnatal results were similar. Among the women who screened positive antenatally, the additional question about the need for help had a sensitivity of 58% (95% CI 38%-76%), a specificity of 91% (95% CI 78%-97%), a positive likelihood ratio of 6.86 (95% CI 2.16-21.7) and a negative likelihood ratio of 0.45 (95% CI 0.25-0.80), with lower sensitivity and higher specificity postnatally. INTERPRETATION: Negative responses to both of the case-finding questions showed acceptable accuracy for ruling out perinatal depression. For positive responses, the use of a third question about the need for help improved specificity and the ability to rule in depression.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.039 | 0.126 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".