Testing guidelines during times of crisis: challenges and limitations of developing rapid and living guidelines
Bibliographic record
Abstract
BACKGROUND: The start of the COVID-19 pandemic presented a situation in which there was an urgent need for decision-making that relates to diagnosis, but the evidence was lacking, of low certainty or constantly changing. Rapid and living guideline development methods were needed and had to be applied to rigorous guideline approaches, such as the Grading of Recommendations Assessment, Development, and Evaluation approach. OBJECTIVES: To describe the process of developing rapid diagnosis guidelines when there is limited and imperfect available data at the time of crisis. SOURCES: Case example from four Infectious Disease Society of America COVID-19 diagnostic guidelines. CONTENT: As the world was experiencing panic with COVID-19, there were serious doubts about the feasibility of following a rigorous process for guideline development when timeliness was of extreme value. The Infectious Disease Society of America guideline panels supported by several methodologists strongly believed that at times of crisis, it is more important than ever to follow a rigorous process. The panel adopted a rapid and living systematic review methodology and applied the Grading of Recommendations Assessment, Development and Evaluation approach to four diagnosis guidelines despite the challenges of scarce and dynamic evidence. We describe the methodological details of the rapid and living approach (data extraction, meta-analysis, Evidence to Decision framework, and recommendation development), the challenge of resources, the challenge of scarce evidence, the challenge of rapidly changing evidence, as well as 'wins' from the Infectious Disease Society of America experience. IMPLICATIONS: Mitigation of pandemics relies on rapid and accurate diagnosis, which is challenged by many knowledge gaps. This necessitates emerging evidence is rapidly incorporated in a living fashion with several decisional and contextual factors to ensure the best public health strategies and care for patients. This process must be systematic and transparent for developing trustworthy guidelines and should be supported by all stakeholders, including researchers, editors, publishers, professional societies, and policymakers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.090 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".