MétaCan
Menu
Back to cohort

Data-Driven Cutoff Selection for the Patient Health Questionnaire-9 Depression Screening Tool

2024· article· en· W4404634668 on OpenAlexafffund
Brooke Levis, Parash Mani Bhandari, Dipika Neupane, Suiqiong Fan, Ying Sun, Chen He, Yin Wu, Ankur Krishnan, Zelalem Negeri, Mahrukh Imran, Danielle B. Rice, Kira E. Riehm, Marleine Azar, A.H. Levis, Jill Boruff, Pim Cuijpers, Simon Gilbody, John P. A. Ioannidis, Lorie A. Kloda, Scott B. Patten, Roy C. Ziegelstein, Daphna Harel, Yemisi Takwoingi, Sarah Markham, Sultan H. Alamri, Dagmar Amtmann, Bruce Arroll, Liat Ayalon, Hamid Reza Baradaran, Anna Beraldi, Charles N. Bernstein, Arvin Bhana, Charles H. Bombardier, Ryna Imma Buji, Peter Butterworth, Gregory Carter, Marcos Hortes Nisihara Chagas, Juliana C.N. Chan, Lai Fong Chan, Dixon Chibanda, Kerrie Clover, Aaron Conway, Yeates Conwell, Federico M. Daray, Janneke M. de Man‐van Ginkel, Jesse R. Fann, Felix Fischer, Sally Field, Jane Fisher, Daniel Fung, Bizu Gelaye, Leila Gholizadeh, Felicity Goodyear‐Smith, Eric Green, Catherine G. Greeno, Brian J. Hall, Liisa Hantsoo, Martin Härter, Leanne Hides, Stevan E. Hobfoll, Simone Honikman, Thomas Hyphantis, Masatoshi Inagaki, María Iglesias-González, Hong Jin Jeon, Nathalie Jetté, Mohammad E. Khamseh, Kim M. Kiely, Brandon A. Kohrt, Yunxin Kwan, Ma. Asunción Lara, Holly Frances Levin-Aspenson, Shen‐Ing Liu, Manote Lotrakul, Sônia Regina Loureiro, Bernd Löwe, Nagendra P. Luitel, Crick Lund, Ruth Ann Marrie, Laura Marsh, Brian P. Marx, Anthony McGuire, Sherina Mohd Sidik, Tiago N. Munhoz, Kumiko Muramatsu, Juliet Nakku, Laura Navarrete, Flávia L. Osório, Brian W. Pence, Philippe Persoons, Inge Petersen, Angelo Picardi, Stephanie L. Pugh, Terence J. Quinn, Elmārs Rancāns, Sujit D. Rathod, Katrin Reuter, Alasdair G Rooney, Iná S. Santos, Miranda T. Schram, Juwita Shaaban, Eileen H. Shinn, Abbey Sidebottom, Adam Simning, Lena Spangenberg, Lesley Stafford, Sharon C. Sung, Keiko Suzuki, Pei Lin Lynnette Tan, Martin Taylor‐Rowan, Thach Tran, Alyna Turner, Christina M. van der Feltz‐Cornelis, Thandi van Heyningen, Paul A. Vöhringer, Lynne I. Wagner, JianLi Wang, David Watson, Jennifer White, Mary A. Whooley, Kirsty Winkley, Karen Wynter, Mitsuhiko Yamada, Qing Zhi Zeng, Yuying Zhang, Brett D. Thombs, Andrea Benedetti

Bibliographic record

VenueJAMA Network Open · 2024
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsDalhousie UniversityUniversity of ManitobaUniversity of CalgaryMcGill UniversityMcMaster UniversityMcGill University Health CentreUniversity of WaterlooJewish General Hospital
FundersMenzies Centre for Australian Studies, King's College London, University of LondonSamsungCanadian Institutes of Health ResearchUniversity of IoanninaBerlin Institute of HealthUniversity of Cape TownHumboldt-Universität zu BerlinUniversidad de Buenos AiresChinese University of Hong KongInyuvesi Yakwazulu-NataliNational Cancer InstituteNew York University ShanghaiKing Abdulaziz UniversityMedical Center, University of RochesterBar-Ilan UniversityQueensland University of TechnologyUniversitätsklinikum Hamburg-EppendorfFreie Universität BerlinMonash UniversityAustralian National UniversityUniversidade de São PauloShimane UniversityUniversiti Kebangsaan MalaysiaGeorge Washington UniversityKing's College LondonMcGill University Health CentreMcGill UniversityUniversity of Technology SydneyHarvard T.H. Chan School of Public HealthInstituto Nacional de Psiquiatría Ramón de la Fuente MuñizSungkyunkwan UniversityYork UniversityUniversity of RochesterUniversity of WashingtonJohns Hopkins UniversityUniversity of PittsburghVrije Universiteit AmsterdamUniversity of QueenslandUniversity of WollongongDuke Global Health Institute, Duke UniversityIran University of Medical Sciences
KeywordsCutoffYouden's J statisticPopulationPatient Health QuestionnaireDepression (economics)MedicineStatisticsReceiver operating characteristicMathematicsDepressive symptomsPsychiatry

Abstract

fetched live from OpenAlex

Importance: Test accuracy studies often use small datasets to simultaneously select an optimal cutoff score that maximizes test accuracy and generate accuracy estimates. Objective: To evaluate the degree to which using data-driven methods to simultaneously select an optimal Patient Health Questionnaire-9 (PHQ-9) cutoff score and estimate accuracy yields (1) optimal cutoff scores that differ from the population-level optimal cutoff score and (2) biased accuracy estimates. Design, Setting, and Participants: This study used cross-sectional data from an existing individual participant data meta-analysis (IPDMA) database on PHQ-9 screening accuracy to represent a hypothetical population. Studies in the IPDMA database compared participant PHQ-9 scores with a major depression classification. From the IPDMA population, 1000 studies of 100, 200, 500, and 1000 participants each were resampled. Main Outcomes and Measures: For the full IPDMA population and each simulated study, an optimal cutoff score was selected by maximizing the Youden index. Accuracy estimates for optimal cutoff scores in simulated studies were compared with accuracy in the full population. Results: The IPDMA database included 100 primary studies with 44 503 participants (4541 [10%] cases of major depression). The population-level optimal cutoff score was 8 or higher. Optimal cutoff scores in simulated studies ranged from 2 or higher to 21 or higher in samples of 100 participants and 5 or higher to 11 or higher in samples of 1000 participants. The percentage of simulated studies that identified the true optimal cutoff score of 8 or higher was 17% for samples of 100 participants and 33% for samples of 1000 participants. Compared with estimates for a cutoff score of 8 or higher in the population, sensitivity was overestimated by 6.4 (95% CI, 5.7-7.1) percentage points in samples of 100 participants, 4.9 (95% CI, 4.3-5.5) percentage points in samples of 200 participants, 2.2 (95% CI, 1.8-2.6) percentage points in samples of 500 participants, and 1.8 (95% CI, 1.5-2.1) percentage points in samples of 1000 participants. Specificity was within 1 percentage point across sample sizes. Conclusions and Relevance: This study of cross-sectional data found that optimal cutoff scores and accuracy estimates differed substantially from population values when data-driven methods were used to simultaneously identify an optimal cutoff score and estimate accuracy. Users of diagnostic accuracy evidence should evaluate studies of accuracy with caution and ensure that cutoff score recommendations are based on adequately powered research or well-conducted meta-analyses.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.103
metaresearch head score (Gemma)0.008
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Scholarly communication, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.910
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.1030.008
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0010.000
Scholarly communication0.0070.001
Open science0.0040.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.659
GPT teacher head0.536
Teacher spread0.124 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations21
Published2024
Admission routes2
Has abstractyes

Explore more

Same venueJAMA Network OpenSame topicMeta-analysis and systematic reviewsFrench-language works237,207