MétaCan
Menu
Back to cohort
Record W2342283369 · doi:10.1148/radiol.2016152229

Meta-Analyses of Diagnostic Accuracy in Imaging Journals: Analysis of Pooling Techniques and Their Effect on Summary Estimates of Diagnostic Accuracy

2016· review· en· W2342283369 on OpenAlexafffund
Trevor A. McGrath, Matthew D. F. McInnes, Daniël A. Korevaar, Patrick M. Bossuyt

Bibliographic record

VenueRadiology · 2016
Typereview
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsOttawa HospitalUniversity of Ottawa
FundersUniversity of Ottawa
KeywordsMedicineMeta-analysisBivariate analysisUnivariateConfidence intervalReceiver operating characteristicSubspecialtyPoolingStatisticsRandom effects modelForest plotMEDLINEMedical physicsMultivariate statisticsInternal medicinePathologyArtificial intelligenceMathematicsComputer science

Abstract

fetched live from OpenAlex

Purpose To determine whether authors of systematic reviews of diagnostic accuracy studies published in imaging journals used recommended methods for meta-analysis, and to evaluate the effect of traditional methods on summary estimates of sensitivity and specificity. Materials and Methods Medline was searched for published systematic reviews that included meta-analysis of test accuracy data limited to imaging journals published from January 2005 to May 2015. Two reviewers independently extracted study data and classified methods for meta-analysis as traditional (univariate fixed- or random-effects pooling or summary receiver operating characteristic curve) or recommended (bivariate model or hierarchic summary receiver operating characteristic curve). Use of methods was analyzed for variation with time, geographical location, subspecialty, and journal. Results from reviews in which study authors used traditional univariate pooling methods were recalculated with a bivariate model. Results Three hundred reviews met the inclusion criteria, and in 118 (39%) of those, authors used recommended meta-analysis methods. No change in the method used was observed with time (r = 0.54, P = .09); however, there was geographic (χ(2) = 15.7, P = .001), subspecialty (χ(2) = 46.7, P < .001), and journal (χ(2) = 27.6, P < .001) heterogeneity. Fifty-one univariate random-effects meta-analyses were reanalyzed with the bivariate model; the average change in the summary estimate was -1.4% (P < .001) for sensitivity and -2.5% (P < .001) for specificity. The average change in width of the confidence interval was 7.7% (P < .001) for sensitivity and 9.9% (P ≤ .001) for specificity. Conclusion Recommended methods for meta-analysis of diagnostic accuracy in imaging journals are used in a minority of reviews; this has not changed significantly with time. Traditional (univariate) methods allow overestimation of diagnostic accuracy and provide narrower confidence intervals than do recommended (bivariate) methods. (©) RSNA, 2016 Online supplemental material is available for this article.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.092
metaresearch head score (Gemma)0.522
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Meta-epidemiology (broad), Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Meta-epidemiology (broad)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.938
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0920.522
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0420.013
Bibliometrics0.0060.005
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0030.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.679
GPT teacher head0.582
Teacher spread0.096 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designOther design
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations57
Published2016
Admission routes2
Has abstractyes

Explore more

Same venueRadiologySame topicMeta-analysis and systematic reviewsFrench-language works237,207