MétaCan
Menu
Back to cohort
Record W4385742310 · doi:10.2196/48444

Patient Health Questionnaire-9 Item Pairing Predictiveness for Prescreening Depressive Symptomatology: Machine Learning Analysis

2023· article· en· W4385742310 on OpenAlexvenueno aff
Darragh Glavin, Eoin Martino Grua, Carina Akemi Nakamura, Márcia Scazufca, Edinilza Ribeiro dos Santos, Gloria Hoi Yan Wong, William Hollingworth, T. J. Peters, Ricardo Araya, Pepijn Van de Ven

Bibliographic record

VenueJMIR Mental Health · 2023
Typearticle
Languageen
FieldDecision Sciences
TopicPsychometric Methodologies and Testing
Canadian institutionsnot available
FundersMedical Research CouncilConselho Nacional de Desenvolvimento Científico e TecnológicoFundação de Amparo à Pesquisa do Estado de São PauloScience Foundation IrelandDepartment of Health and Social CareWellcome Trust
KeywordsPatient Health QuestionnaireMajor depressive disorderMoodAnhedoniaItem response theoryPsychologyClinical psychologyLogistic regressionGeneralizability theoryOrdered logitMedicineArtificial intelligencePsychiatryMachine learningPsychometricsDepressive symptomsComputer scienceDevelopmental psychologyCognition

Abstract

fetched live from OpenAlex

BACKGROUND: Anhedonia and depressed mood are considered the cardinal symptoms of major depressive disorder. These are the first 2 items of the Patient Health Questionnaire (PHQ)-9 and comprise the ultrabrief PHQ-2 used for prescreening depressive symptomatology. The prescreening performance of alternative PHQ-9 item pairings is rarely compared with that of the PHQ-2. OBJECTIVE: This study aims to use machine learning (ML) with the PHQ-9 items to identify and validate the most predictive 2-item depressive symptomatology ultrabrief questionnaire and to test the generalizability of the best pairings found on the primary data set, with 6 external data sets from different populations to validate their use as prescreening instruments. METHODS: All 36 possible PHQ-9 item pairings (each yielding scores of 0-6) were investigated using ML-based methods with logistic regression models. Their performances were evaluated based on the classification of depressive symptomatology, defined as PHQ-9 scores ≥10. This gave each pairing an equal opportunity and avoided any bias in item pairing selection. RESULTS: The ML-based PHQ-9 items 2 and 4 (phq2&4), the depressed mood and low-energy item pairing, and PHQ-9 items 2 and 8 (phq2&8), the depressed mood and psychomotor retardation or agitation item pairing, were found to be the best on the primary data set training split. They generalized well on the primary data set test split with area under the curves (AUCs) of 0.954 and 0.946, respectively, compared with an AUC of 0.942 for the PHQ-2. The phq2&4 had a higher AUC than the PHQ-2 on all 6 external data sets, and the phq2&8 had a higher AUC than the PHQ-2 on 3 data sets. The phq2&4 had the highest Youden index (an unweighted average of sensitivity and specificity) on 2 external data sets, and the phq2&8 had the highest Youden index on another 2. The PHQ-2≥2 cutoff also had the highest Youden index on 2 external data sets, joint highest with the phq2&4 on 1, but its performance fluctuated the most. The PHQ-2≥3 cutoff had the highest Youden index on 1 external data set. The sensitivity and specificity achieved by the phq2&4 and phq2&8 were more evenly balanced than the PHQ-2≥2 and ≥3 cutoffs. CONCLUSIONS: The PHQ-2 did not prove to be a more effective prescreening instrument when compared with other PHQ-9 item pairings. Evaluating all item pairings showed that, compared with alternative partner items, the anhedonia item underperformed alongside the depressed mood item. This suggests that the inclusion of anhedonia as a core symptom of depression and its presence in ultrabrief questionnaires may be incompatible with the empirical evidence. The use of the PHQ-2 to prescreen for depressive symptomatology could result in a greater number of misclassifications than alternative item pairings.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.009
metaresearch head score (Gemma)0.024
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.009
Threshold uncertainty score0.049

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0090.024
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.198
GPT teacher head0.483
Teacher spread0.285 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations10
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Mental HealthSame topicPsychometric Methodologies and TestingFrench-language works237,207