Exploring variation in the use of feedback from national clinical audits: a realist investigation
Bibliographic record
Abstract
BACKGROUND: National Clinical Audits (NCAs) are a well-established quality improvement strategy used in healthcare settings. Significant resources, including clinicians' time, are invested in participating in NCAs, yet there is variation in the extent to which the resulting feedback stimulates quality improvement. The aim of this study was to explore the reasons behind this variation. METHODS: We used realist evaluation to interrogate how context shapes the mechanisms through which NCAs work (or not) to stimulate quality improvement. Fifty-four interviews were conducted with doctors, nurses, audit clerks and other staff working with NCAs across five healthcare providers in England. In line with realist principles we scrutinised the data to identify how and why providers responded to NCA feedback (mechanisms), the circumstances that supported or constrained provider responses (context), and what happened as a result of the interactions between mechanisms and context (outcomes). We summarised our findings as Context+Mechanism = Outcome configurations. RESULTS: We identified five mechanisms that explained provider interactions with NCA feedback: reputation, professionalism, competition, incentives, and professional development. Professionalism and incentives underpinned most frequent interaction with feedback, providing opportunities to stimulate quality improvement. Feedback was used routinely in these ways where it was generated from data stored in local databases before upload to NCA suppliers. Local databases enabled staff to access data easily, customise feedback and, importantly, the data were trusted as accurate, due to the skills and experience of staff supporting audit participation. Feedback produced by NCA suppliers, which included national comparator data, was used in a more limited capacity across providers. Challenges accessing supplier data in a timely way and concerns about the quality of data submitted across providers were reported to constrain use of this mode of feedback. CONCLUSION: The findings suggest that there are a number of mechanisms that underpin healthcare providers' interactions with NCA feedback. However, there is variation in the mode, frequency and impact of these interactions. Feedback was used most routinely, providing opportunities to stimulate quality improvement, within clinical services resourced to collect accurate data and to maintain local databases from which feedback could be customised for the needs of the service.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".