Impact of Changes in Clinical Practice Guidelines on Assessment of Quality of Care
Bibliographic record
Abstract
BACKGROUND: Measures for pay-for-performance and public reporting programs may be based on clinical practice guidelines. The impact of guideline changes over time-and whether evolving clinical evidence can render measures based on prior guidelines misleading-is not known. OBJECTIVE: To assess the impact of using different percutaneous coronary intervention (PCI) guidelines when evaluating whether PCI was indicated. RESEARCH DESIGN: PCIs from the National Cardiovascular Data Registry's CathPCI registry performed in 2003-2004 were categorized into indication classes (Class I, IIa, IIb, III), using 2001 American College of Cardiology/American Heart Association guidelines for PCI, the guidelines available at the time of the procedures. The same procedures were recategorized using 2005 guidelines, which reflect the best evidence available to clinicians at the time of PCI. Procedures unable to be categorized were labeled as "Not Certain." SUBJECTS: Patients undergoing PCI for stable or unstable angina in 394 hospitals. MEASURES: Number of procedures changing classification categories using 2001 versus 2005 guidelines. RESULTS: A total of 345,779 PCIs were evaluated. Applying 2001 guidelines, 47.9% had Class I indications; 33.3% Class IIa; 5.9% Class IIb; 3.7% Class III; and 9.2% Not Certain. Applying 2005 guidelines to the same procedures, 25.1% had Class I indications; 57.5% Class IIa; 5.5% Class IIb; 3.7% Class III; and 8.3% Not Certain; 41.1% of procedures changed the classification overall. CONCLUSIONS: The changes in guidelines resulted in a marked shift in whether PCIs done in 2003-2004 were considered indicated. Guideline-based performance measures should be carefully evaluated before implementation to avoid incorrect assessments of quality of care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.170 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".