MétaCan
Menu
Back to cohort
Record W2167473079 · doi:10.1093/humrep/des342

Methodological quality of systematic reviews in subfertility: a comparison of Cochrane and non-Cochrane systematic reviews in assisted reproductive technologies

2012· article· en· W2167473079 on OpenAlexaff
Bethany Windsor, Ivor Popovich, Vanessa Jordan, Marian Showell, Beverley Shea, Cindy Farquhar

Bibliographic record

VenueHuman Reproduction · 2012
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsUniversity of Ottawa
FundersUniversity of Auckland
KeywordsSystematic reviewData extractionCochrane LibraryMedicineMEDLINEMeta-analysisBiologyPathology

Abstract

fetched live from OpenAlex

STUDY QUESTION: Are there differences in the methodological quality of Cochrane systematic reviews (CRs) and non-Cochrane systematic reviews (NCRs) of assisted reproductive technologies? SUMMARY ANSWER: CRs on assisted reproduction are of higher methodological quality than similar reviews published in other journals. WHAT IS KNOWN ALREADY: The quality of systematic reviews varies. STUDY DESIGN, SIZE AND DURATION: This was a cross-sectional study of 30 CR and 30 NCR systematic reviews that were randomly selected from the eligible reviews identified from a literature search for the years 2007-2011. MATERIALS, SETTING AND METHODS: We extracted data on the reporting and methodological characteristics of the included systematic reviews. We assessed the methodological quality of the reviews using the 11-domain Measurement Tool to Assess the Methodological Quality of Systematic Reviews (AMSTAR) tool and subsequently compared CR and NCR systematic reviews. MAIN RESULTS AND THE ROLE OF CHANCE: The AMSTAR quality assessment found that CRs were superior to NCRs. For 10 of 11 AMSTAR domains, the requirements were met in >50% of CRs, but only 4 of 11 domains showed requirements being met in >50% of NCRs. The strengths of CRs are the a priori study design, comprehensive literature search, explicit lists of included and excluded studies and assessments of internal validity. Significant failings in the CRs were found in duplicate study selection and data extraction (67% meeting requirements), assessment for publication bias (53% meeting requirements) and reporting of conflicts of interest (47% meeting requirements). NCRs were more likely to contain methodological weaknesses as the majority of the domains showed <40% of reviews meeting requirements, e.g. a priori study design (17%), duplicate study selection and data extraction (17%), assessment of study quality (27%), study quality in the formulation of conclusions (23%) and reporting of conflict of interests (10%). LIMITATIONS, REASONS FOR CAUTION: The AMSTAR assessment can only judge what is reported by authors. Although two of the five authors are involved in the production of CRs, the risk of bias was reduced by not involving these authors in the assessment of the systematic review quality. WIDER IMPLICATIONS OF THE FINDINGS: Not all systematic reviews are equal. The reader needs to consider the quality of the systematic review when they consider the results and the conclusions of a systematic review. STUDY FUNDING/COMPETING INTEREST(S): There are no conflicts with any commercial organization. Funding was provided for the students by the summer studentship programme of the Faculty of Medical and Health Sciences of the University of Auckland.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Direct model labels (unvalidated)

Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.

Model armCategoriesStudy designConfidence
gemmaMetaresearch
Domain: Methods · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Observationallow
gptMetaresearchMeta-epidemiology (broad)
Domain: Methods · Genre: Review
About the Canadian research system: no · About a Canadian topic: no
Systematic reviewhigh
models splitAgreement compares identical category sets and study designs across arms.

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.443
metaresearch head score (Gemma)0.768
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (broad)
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.976
Threshold uncertainty score0.687

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4430.768
Meta-epidemiology (narrow)0.0030.004
Meta-epidemiology (broad)0.0240.037
Bibliometrics0.0430.041
Science and technology studies0.0030.006
Scholarly communication0.0150.012
Open science0.0050.010
Research integrity0.0050.005
Insufficient payload (model declined to judge)0.0040.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.908
GPT teacher head0.626
Teacher spread0.282 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Labeled directly by 2 models reading the full record.

MetaresearchMeta-epidemiology (broad)

The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.

Study designObservational · Systematic review
DomainMethods
GenreEmpirical · Review

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations50
Published2012
Admission routes1
Has abstractyes

Explore more

Same venueHuman ReproductionSame topicMeta-analysis and systematic reviewsCategoryMetaresearchFrench-language works237,207