MétaCan
Menu
Back to cohort
Record W4321163600 · doi:10.1097/js9.0000000000000114

Should the reporting certainty of evidence for meta-analysis of observational studies using GRADE be revisited?

2023· article· en· W4321163600 on OpenAlexaboutno aff
Mahmoud Yousefifard, Arman Shafiee

Bibliographic record

VenueInternational Journal of Surgery · 2023
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsnot available
Fundersnot available
KeywordsObservational studyGrading (engineering)Systematic reviewMedicineCertaintyGuidelineMeta-analysisEvidence-based medicineMEDLINEPsychological interventionCritical appraisalMedical physicsAlternative medicinePathologyPsychiatry

Abstract

fetched live from OpenAlex

Highlights Assessing the certainty of results is an inevitable part of systematic reviews. Grading of Recommendations Assessment, Development, and Evaluation is a widely used tool for assessing certainty of evidence. Using Grading of Recommendations Assessment, Development, and Evaluation for observational studies meta-analyses is accompanied by limitations. Dear Editor, It is evident in the literature that aside from reporting the synthesis from the predefined questions, one must be aware of the validity and reliability of their estimates. Reporting the level of evidence will guide policymakers and practitioners to apply evidence in routine practice. In 2004, The Grading of Recommendations Assessment, Development, and Evaluation (GRADE) working group developed an approach for systematic reviews to provide transparent information on the certainty of evidence. The approach was initially designed to address the effectiveness of interventions in systematic reviews. However, it has been widely used by researchers as an evidence-ranking scheme to assess the certainty of the evidence of all types of systematic reviews, including prognostic, diagnostic, and prevalence studies. In our experience, the journal reviewers and editors suggested reporting the certainty of the evidence for prognostic/diagnostic systematic reviews using the GRADE guideline. However, applying the GRADE guideline for the rating of the prognostic/diagnostic systematic reviews have important limitations. Here, we will address some issues regarding using this approach in systematic reviews and meta-analyses of observational studies. The GRADE approach has four levels of evidence: very low, low, moderate, and high. Eight GRADE criteria must be evaluated to report the certainty of each synthesized outcome. Among them, five decreases the grade of evidence (risk of bias, imprecision, inconsistency, indirectness, and publication bias)1. In the case of methodologically robust observational research, GRADE advises raising the quality of the evidence (large magnitude of effect, dose-response gradient, and all plausible residual confounding). Most of the published systematic reviews have stated they reported their study in line with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guideline2. According to PRISMA guideline, authors must specify how they assess certainty (or confidence) in the body of evidence for an outcome. The GRADE approach suggests evidence from randomized clinical trials must begin with a high level, while evidence based on observational studies starts with a low level. Therefore, it seems that researchers of meta-analyses of observational studies might not be as eager to utilize this method to assess the certainty of evidence because the quality of a meta-analysis of an observational study will be mostly judged as low or very low. In addition, in the latest revision of GRADE guideline, the risk of bias assessment may cause rating down of evidence three levels3. For example, O’Keeffe et al.4 published a meta-analysis on the relation of smoking and risk of lung cancer in 2018. In this well-written systematic review, a high risk of bias was observed according to the Newcastle–Ottawa scale, and there was also considerable heterogeneity (89.4–98.8%). According to the GRADE guideline, the certainty of the evidence is rated down five levels: three points due to an extremely serious risk of bias and two points due to very serious heterogeneity. As well, we can rate up the level of evidence two to three levels since there was a large magnitude of effect and potential dose-response gradients (increasing the risk of cancer with increasing the number of cigarettes per day). As a conclusion, the overall certainty of evidence for the relationship between smoking and lung cancer derived from the O’Keeffe and colleagues study is ‘low to very low.’ Therefore, based on the definitions given by the GRADE, this relationship will probably be markedly different from the actual estimated effect. While there is a global consensus on the independent hazardous effect of smoking on the increasing risk of lung cancer. Although randomized clinical trials provide a higher level of evidence than observational studies to assess the safety and efficacy of intervention, in the assessment of the diagnostic/prognostic value of an indicator, observational studies can provide the optimum evidence available5. Furthermore, based on the Centre for Evidence-Based Medicine (CEBM) recommendation, systematic reviews of observational studies evaluated to address prognosis, diagnosis, and prevalence of a predefined question are among the studies with the highest levels of evidence. In conclusion, it seems that the prespecified low rating of the quality of evidence derived from observational studies proposed by the GRADE may lead to an excessive decrease in the quality of evidence, and most of the authors of meta-analyses of observational studies tend not to report the certainty of evidence throughout their manuscript. In addition, rating down of certainty three levels for application of new risk of bias tools such as ROBINS-I may not applicable for systematic review of observational studies. We propose to define a new approach for rating the level of evidence in such a way that the rating for systematic reviews of observational studies is designed based on the CEBM guideline. For example, the CEBM suggests the level of evidence provided by well-designed cohort studies and cross-sectional studies to investigate the prognostic and diagnostic value of an indicator is ‘high.’ Therefore, if a systematic review of well-designed cross-sectional studies is conducted to assess the diagnostic value of an indicator, it is better to start rating its certainty from high-quality evidence. Moreover, ‘extremely serious’ risk of bias due to using the new risk of bias tools should be apply only in systematic reviews of treatment effect. Ethical approval Not applicable. Sources of funding None. Author contribution M.Y. contributed to conception of the manuscript, drafted the manuscript, circulated for review, and revised the final manuscript. A.S. drafted the manuscript, circulated for review, and revised the final manuscript. Conflict of interest disclosure The authors declare that they have no financial conflict of interest with regard to the content of this report. Research registration unique identifying number (UIN) None. Guarantor Mahmoud Yousefifard and Arman Shafiee. Data statement None.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.236
metaresearch head score (Gemma)0.351
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (broad), Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Meta-analysis · Consensus signal: Meta-analysis
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.140
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.2360.351
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0060.013
Bibliometrics0.0020.003
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.995
GPT teacher head0.715
Teacher spread0.280 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designMeta-analysis
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations21
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueInternational Journal of SurgerySame topicMeta-analysis and systematic reviewsFrench-language works237,207