A25 CRITICAL APPRAISAL OF GI ENDOSCOPY CLINICAL PRACTICE GUIDELINES DURING THE COVID-19 PANDEMIC
Bibliographic record
Abstract
Abstract Background Clinical Practice Guidelines (CPGs) are integral during a pandemic, offering guidance to clinicians through uncertainty. Existing literature has established that the need for rapid publication of CPGs during previous infectious disease outbreaks resulted in less rigorous guidelines. CPGs were rapidly developed since the onset of the pandemic in December 2019, providing guidance in gastrointestinal (GI) endoscopy, an area where COVID-19 may pose risk of transmission. Aims To evaluate the quality of GI endoscopy guidelines developed during the COVID-19 pandemic and to compare these with (a) endoscopy CPGs developed prior to the pandemic; (b) CPGs for other endoscopic topics unrelated to COVID-19; and, (c) non-endoscopic CPGs published during the pandemic. Methods We systematically searched Medline, Embase and Scopus for CPGs published by GI societies from January 1, 2018 to December 31, 2020. A grey literature search was conducted. Two authors screened full-texts. In this interim analysis, CPGs were grouped based on publication year: before 2020, or 2020. Endoscopy CPGs published in 2020 were categorized as COVID or non-COVID related. Two authors independently assessed the CPGs using the AGREE II tool, consisting of six domains for evaluating guidelines. A domain score of 60 was set as a threshold to indicate good quality. Results There were 70 endoscopy guidelines and 27 CPGs focused on other GI topics. The mean overall scores were 69% (±12%) for endoscopy CPGs published before 2020 (n=28), and 51% (±23%) for CPGs published in 2020 (n=42). For individual AGREE II domains, mean scores for pre-2020 CPGs ranged from 33.11 (±17.39) in Applicability to 81.55 (±10.37) in Clarity of Presentation. For CPGs published during COVID-19, mean domain scores ranged from 34.18 (±10.52) in Applicability to 75.26 (±13.85) in Clarity of Presentation. 21 of 42 CPGs published in 2020 were related to COVID. Mean overall scores were 35% (±20%) for COVID-related CPGs and 67% (±13%) for non-COVID-19 CPGs. For COVID-19 CPGs, scores ranged from 27.88 (±20.31) in Rigour of Development to 69.58 (±10.81) in Scope and Purpose. For non-COVID CPGs, the scores ranged from 37.30 (±8.93) in Applicability to 84.52 (±5.93) in Clarity of Presentation. Conclusions The difference in overall scores between COVID-19 endoscopy CPGs and non-COVID endoscopy CPGs may suggest that the urgency to disseminate COVID-19 information decreased CPG quality or completeness of reporting. This interim analysis is limited by the lack of distinction between peer-reviewed CPGs and non-peer reviewed recommendations. Given the importance of CPGs in clinical decision making, it is important to ensure that the rapid development of guidelines does not compromise quality and rigour. Funding Agencies None
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.351 | 0.757 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.007 |
| Bibliometrics | 0.041 | 0.026 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.009 | 0.010 |
| Open science | 0.005 | 0.007 |
| Research integrity | 0.005 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".