Standardized Reporting of Clinical Practice Guidelines: A Proposal from the Conference on Guideline Standardization
Bibliographic record
Abstract
Despite enormous energies invested in authoring clinical practice guidelines, the quality of individual guidelines varies considerably. The Conference on Guideline Standardization (COGS) was convened in April 2002 to define a standard for guideline reporting that would promote guideline quality and facilitate implementation. Twenty-three people with expertise and experience in guideline development, dissemination, and implementation participated. A list of candidate guideline components was assembled from the Institute of Medicine Provisional Instrument for Assessing Clinical Guidelines, the National Guideline Clearinghouse, the Guideline Elements Model, and other published guideline models. In a 2-stage modified Delphi process, panelists first rated their agreement with the statement that "[Item name] is a necessary component of practice guidelines" on a 9-point scale. An individualized report was prepared for each panelist; the report summarized the panelist's rating for each item and the median and dispersion of rankings of all the panelists. In a second round, panelists separately rated necessity for validity and necessity for practical application. Items achieving a median rank of 7 or higher on either scale, with low disagreement index, were retained as necessary guideline components. Representatives of 22 organizations active in guideline development reviewed the proposed items and commented favorably. Closely related items were consolidated into 18 topics to create the COGS checklist. This checklist provides a framework to support more comprehensive documentation of practice guidelines. Most organizations that are active in guideline development found the component items to be comprehensive and to fit within their existing development methods.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.478 | 0.355 |
| Meta-epidemiology (narrow) | 0.004 | 0.005 |
| Meta-epidemiology (broad) | 0.004 | 0.010 |
| Bibliometrics | 0.012 | 0.010 |
| Science and technology studies | 0.009 | 0.015 |
| Scholarly communication | 0.026 | 0.019 |
| Open science | 0.016 | 0.026 |
| Research integrity | 0.074 | 0.081 |
| Insufficient payload (model declined to judge) | 0.005 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".