Evidence on depression screening tool accuracy: an assessment of the appropriateness of primary study samples and the methodological quality and reporting transparency of evidence syntheses
Bibliographic record
Abstract
Background: Depression accounts for more years lived with disability than any other medical condition.Routine screening for depression has been recommended to improve access to depression care.However, a previous study from 2011 found that a large proportion of the diagnostic accuracy of depression screening studies have been conducted in samples that inappropriately include patients currently diagnosed or being treated for depression.These findings from primary studies suggest that estimates of the diagnostic test accuracy of depression screening tools may be exaggerated.Similarly, concerns have been raised regarding exaggerated accuracy estimates based on meta-analyses of the diagnostic test accuracy of depression screening tools.Meta-analyses have the ability to account for biases that may be present within primary studies of diagnostic test accuracy of depression screening tools.However, the extent to which this occurs is currently unknown, thus, the objectives of the three studies in this thesis were to determine (1) the proportion of recent diagnostic test accuracy primary studies of depression screening tools that appropriately excluded patients with a current diagnosis of depression and assess if recent meta-analyses of depression screening tools noted the potential bias due to the inappropriate inclusion of these patients (study 1), (2) the quality and transparency of diagnostic test accuracy meta-analyses of depression screening tools by applying an adapted A Measurement Tool to Assess Systematic Reviews (AMSTAR) quality tool (study 2) and an adapted Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) for Abstracts reporting checklist to a sample of diagnostic test accuracy of depression screening tool meta-analyses (study 3).Methods: To gather recent primary studies of diagnostic accuracy of depression screening tools we searched the PubMed online database from January 1, 2013 through March 27, 2015.We determined the proportion of these studies that accurately excluded patients with a current v diagnosis of depression or were being treated for depression.To obtain recent diagnostic test accuracy meta-analyses of depression screening tools, we searched PubMed and PsycINFO from January 1, 2011 through October 31, 2014 and assessed if meta-analyses mentioned potential bias of inappropriate samples.Next, we obtained all meta-analyses published in PubMed and PsycINFO between January 1, 2005 through October 31, 2014.We evaluated 1) the quality of meta-analyses using an adapted AMSTAR and 2) the transparency of abstract reporting using an adapted PRISMA for Abstract tools.Results: Only 5 of the 90 (5.6%) primary studies specifically excluded already diagnosed or treated patients, while none of 5 eligible meta-analyses commented on the possible bias from including primary studies that used inappropriate samples.The quality of 16 meta-analyses were evaluated using an adapted AMSTAR tool, where five of 14 AMSTAR items were fulfilled by at least half of the meta-analyses.Two of the 16 meta-analyses received a yes rating for 8 and 10 of the 14 items.The abstracts of these 16 meta-analyses were also assessed for the transparency of reporting using an adapted PRISMA for Abstracts tool where only 2 of 14 items were fulfilled by at least half of the meta-analyses.Only one meta-analyses received ratings of yes for more than half of the PRISMA for Abstract items.Conclusions: Current estimates of the accuracy of depression screening tools may still be over exaggerated partially due to the inappropriate sample of patients being studied.Furthermore, current meta-analyses and abstracts of diagnostic test accuracy of depression screening tools may misrepresent results as there is a lack of high quality meta-analyses with transparent reporting.vi RÉSUMÉ Contexte: La dépression cause plus d'années vécues avec une incapacité que toute autre condition médicale.Le dépistage systématique de la dépression a été recommandé afin d'améliorer l'accès aux soins pour la dépression.Cependant, une étude antérieure de 2011 a conclu qu'une grande proportion d'études sur la précision diagnostique de dépistage de dépression a été menée avec des échantillons inappropriés qui inclus des patients actuellement diagnostiqués ou traités pour la dépression.Ces résultats d'études primaires suggèrent que les estimations de la précision de tests diagnostiques d'outils de dépistage de dépression peuvent être exagérées.De même, des préoccupations ont été soulevées au sujet d'estimations de précision exagérées basées sur des méta-analyses de précision de tests diagnostiques d'outils de dépistage de dépression.Les méta-analyses offrent la possibilité de tenir compte de la partialité potentiellement présente dans les études primaires de précision de tests diagnostiques d'outils de dépistage de dépression.Cependant, l'étendue de ce phénomène n'est pas actuellement connue, donc les objectifs des trois études de cette thèse étaient de déterminer (1) la proportion de récentes études primaires sur la précision de tests diagnostiques d'outils de dépistage de dépression qui excluait de manière appropriée les patients avec un diagnostic actuel de dépression et d'évaluer si les méta-analyses récentes sur les outils de dépistage de dépression ont noté la partialité potentielle dûe à l'inclusion inappropriée de ces patients (étude 1), ( 2) la qualité et la transparence des méta-analyses sur la précision de tests diagnostiques d'outils de dépistage de dépression en appliquant l'outil adapté «AMSTAR » (étude 2) et l'outil adapté «PRISMA for abstracts» à un échantillon de méta-analyses sur la précision de tests diagnostiques d'outils de dépistage de dépression (étude 3).Méthodes: Afin de rassembler les récentes études primaires sur la précision de tests diagnostiques d'outils de dépistage de dépression, nous avons recherché dans la base de données en ligne PubMed du 1 er janvier 2013 au 27 mars 2015.Nous avons déterminé la proportion d'études qui excluait vii correctement les patients avec un diagnostic actuel de dépression ou ceux en traitement pour dépression.Pour obtenir les récentes méta-analyses sur la précision de tests diagnostiques d'outils de dépistage de dépression, nous avons effectué une recherche sur PubMed et PsycINFO du 1 er janvier 2011 au 31 octobre 2014 et avons évalué si les méta-analyses mentionnaient une partialité potentiellement dûe à un échantillon inapproprié.Ensuite, nous avons obtenu toutes les métaanalyses publiées dans PubMed et
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.765 | 0.929 |
| Meta-epidemiology (narrow) | 0.003 | 0.006 |
| Meta-epidemiology (broad) | 0.012 | 0.024 |
| Bibliometrics | 0.025 | 0.026 |
| Science and technology studies | 0.004 | 0.012 |
| Scholarly communication | 0.016 | 0.010 |
| Open science | 0.007 | 0.010 |
| Research integrity | 0.008 | 0.007 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".